WO2023152803A1 - 音声認識装置、及びコンピュータが読み取り可能な記録媒体 - Google Patents
音声認識装置、及びコンピュータが読み取り可能な記録媒体 Download PDFInfo
- Publication number
- WO2023152803A1 WO2023152803A1 PCT/JP2022/004938 JP2022004938W WO2023152803A1 WO 2023152803 A1 WO2023152803 A1 WO 2023152803A1 JP 2022004938 W JP2022004938 W JP 2022004938W WO 2023152803 A1 WO2023152803 A1 WO 2023152803A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- adjusted
- voice
- speech recognition
- generation unit
- recognition device
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/20—Speech recognition techniques specially adapted for robustness in adverse environments, e.g. in noise, of stress induced speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/26—Speech to text systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
- G10L15/32—Multiple recognisers used in sequence or in parallel; Score combination systems therefor, e.g. voting systems
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/223—Execution procedure of a spoken command
Definitions
- the present invention relates to a speech recognition device and a computer-readable recording medium.
- the operation part of the device has many buttons and operation screens, but the operation is complicated and it takes time to master.
- a voice input interface can execute a desired operation simply by speaking a voice command. Therefore, attempts have been made to improve operability using a voice input interface.
- the voice commands used to operate the device can be assumed depending on the type of device that uses the voice command, the site where the device is installed, and the operation details of the device. Therefore, expected voice commands can be created in grammar (syntax and words). See Patent Document 1, for example.
- the speech recognition apparatus generates a plurality of speech signals obtained by finely adjusting predetermined attributes (waveform parameters) of the input speech signal, and treats each of them as the target of speech recognition. Then, the above problem is solved by using the mode of the recognition result as the correct recognition result.
- one aspect of the present disclosure is a speech recognition device that recognizes a speech signal input at a manufacturing site and uses it as a speech command, wherein a plurality of different adjustments are made to a predetermined attribute of the input speech signal. and an adjusted waveform group generation unit that generates a plurality of adjusted audio signals corresponding to the adjusted waveform group generation unit, and a speech recognition that performs speech recognition on the audio signal and the plurality of adjusted audio signals output by the adjusted waveform group generation unit and wherein the adjustment performed by the adjusted waveform group generation unit includes speech rate as an attribute to be adjusted.
- Another aspect of the present disclosure is a computer-readable recording medium for recording a program executed by a speech recognition device that recognizes speech signals input at a manufacturing site and uses them as speech commands, an adjusted waveform group generation unit that performs a plurality of different adjustments to a predetermined attribute including an utterance rate of an audio signal and generates a plurality of adjusted audio signals corresponding thereto; and the audio output by the adjusted waveform group generation unit.
- a computer-readable recording medium for recording a program that causes a computer to function as a speech recognition unit that performs speech recognition on a signal and a plurality of the adjusted speech signals.
- the processing accuracy of speech recognition is robust even if a predetermined attribute of the speech waveform is disturbed. Therefore, it is expected that the accuracy rate of voice recognition will also be improved.
- FIG. 1 is a schematic hardware configuration diagram of a speech recognition device according to an embodiment of the present invention
- FIG. 1 is a block diagram showing schematic functions of a speech recognition device according to an embodiment of the present invention
- FIG. It is an example of an adjustment method information registration screen. It is an example of an aggregation method information registration screen.
- FIG. 10 is a diagram showing an example of tallying the mode values of transcribed character strings
- FIG. 10 is a diagram showing an example of counting the median reliability of transcribed character strings.
- FIG. 4 is a block diagram showing schematic functions of a speech recognition device according to another embodiment of the present invention.
- FIG. 1 is a schematic hardware configuration diagram showing essential parts of a speech recognition apparatus according to an embodiment of the present invention.
- the speech recognition device 1 can be mounted on a control device that controls industrial machines 2 installed at a manufacturing site such as a factory.
- the speech recognition device 1 can be implemented on a computer such as a personal computer attached to the control device, a fog computer 6 connected to the control device via a wired or wireless network, a cloud server 7, or the like.
- the voice recognition device 1 according to the present embodiment will be described based on an example in which the voice recognition device 1 is mounted on a control device that controls the industrial machine 2 .
- the CPU 11 included in the speech recognition device 1 is a processor that controls the speech recognition device 1 as a whole.
- the CPU 11 reads the system program stored in the ROM 12 via the bus 22 and controls the entire speech recognition apparatus 1 according to the system program.
- the RAM 13 temporarily stores calculation data, display data, various data input from the outside, and the like.
- the non-volatile memory 14 is composed of, for example, a battery-backed memory (not shown) or an SSD (Solid State Drive), etc., and retains the memory state even when the voice recognition apparatus 1 is powered off.
- the nonvolatile memory 14 stores data acquired from the industrial machine 2, control programs and data read from the external device 72 via the interface 15, control programs and data input via the input device 71, network Control programs and data acquired from other devices via 5 are stored.
- the control program and data stored in the nonvolatile memory 14 may be developed in the RAM 13 at the time of execution/use.
- Various system programs such as a well-known analysis program are pre-written in the ROM 12 .
- the interface 15 is an interface for connecting the CPU 11 of the speech recognition device 1 and an external device 72 such as a USB device. From the external device 72 side, for example, control programs and setting data used for controlling the industrial machine 2 are read. Control programs and setting data edited in the speech recognition apparatus 1 can be stored in the external storage means via the external device 72 .
- a PLC (programmable logic controller) 16 executes a ladder program to control the industrial machine 2 and peripheral devices of the industrial machine 2 (for example, a tool changer, an actuator such as a robot, and a temperature sensor attached to the industrial machine 2). and a plurality of sensors 3) such as a humidity sensor, etc., through the I/O unit 19 to control them. It also receives signals from various switches on an operation panel provided on the main body of the industrial machine 2 and signals from peripheral devices, and passes the signals to the CPU 11 after performing necessary signal processing.
- the interface 20 is an interface for connecting the CPU of the speech recognition device 1 and the wired or wireless network 5 .
- Other industrial machines 4 such as machine tools and electric discharge machines, a fog computer 6, a cloud server 7, and the like are connected to the network 5, and exchange data with the speech recognition apparatus 1 mutually.
- each data read into the memory, data obtained as a result of executing the program, etc. are output via the interface 17 and displayed.
- An input device 71 composed of a keyboard, a pointing device, etc., transfers commands, data, etc. based on operations by an operator to the CPU 11 via the interface 18 .
- the interface 21 is an interface for connecting the CPU 11 of the speech recognition device 1 and the speech sensor 73 .
- the audio sensor 73 may be, for example, a sound collecting device such as a microphone.
- the voice sensor 73 may be attached to, for example, the input device 71, a machine operation panel (not shown), a pendant (portable machine operation panel), or the like. The worker's voice detected by the voice sensor 73 is transferred to the CPU 11 as a voice signal.
- the axis control circuit 30 for controlling the axes provided in the industrial machine 2 receives the axis movement command amount from the CPU 11 and outputs the axis command to the servo amplifier 40 .
- the servo amplifier 40 receives this command and drives the servo motor 50 that moves the axis of the machine tool.
- the axis servomotor 50 incorporates a position/velocity detector, and feeds back a position/velocity feedback signal from this position/velocity detector to the axis control circuit 30 to perform position/velocity feedback control.
- Only one axis control circuit 30, one servo amplifier 40, and one servo motor 50 are shown in the hardware configuration diagram of FIG. only available.
- FIG. 2 is a schematic block diagram of the functions of the speech recognition device 1 according to one embodiment of the present invention. Each function provided in the speech recognition apparatus 1 according to the present embodiment is realized by the CPU 11 provided in the speech recognition apparatus 1 shown in FIG. .
- the speech recognition apparatus 1 of this embodiment includes a speech signal acquisition unit 100, an adjustment method registration unit 110, an adjustment waveform group generation unit 120, a speech recognition unit 130, an aggregation method registration unit 140, an aggregation result generation unit 150, and a command processing unit 160. , and an output unit 170 . Further, in the RAM 13 to the non-volatile memory 14 of the speech recognition apparatus 1, the adjustment method storage unit 180, which is an area for storing the adjustment method data registered by the adjustment method registration unit 110, and the aggregation method registration unit 140 registered A tabulation method storage unit 190, which is an area for storing tabulation method data, is prepared in advance.
- the audio signal acquisition unit 100 acquires the audio signal detected by the audio sensor 73 . Then, an audio signal recognized as one utterance is extracted from the acquired audio signal.
- the audio signal detected by the audio sensor 73 is mainly based on the voice uttered by the operator.
- the voice signal acquisition unit 100 may extract a voice signal corresponding to one utterance of the worker from among them. For example, a state in which the audio signal is equal to or lower than a predetermined level Lv th continues for a predetermined period of time Ts th or more.
- the above audio signal may be cut out as an audio signal corresponding to one utterance. Also, other known audio signal analysis techniques may be used to cut out the audio.
- the audio signal cut out by the audio signal acquisition unit 100 is output to the adjusted waveform group generation unit 120 .
- the adjustment method registration unit 110 receives information on the adjustment method of the speech waveform and registers it in the adjustment method storage unit 180 .
- the information related to the adjustment method includes information related to the attribute of the audio signal to be adjusted and information related to the adjustment range for the attribute. Attributes to be adjusted include, for example, speech rate, amplitude, pitch, formant, SN ratio, and the like.
- the adjustment method registration unit 110 receives, for example, whether or not each attribute is to be adjusted, and if so, what adjustment range to adjust. Then, the received input is used as information related to the adjustment method.
- As the information related to the adjustment width instead of a fixed value, a random number having a maximum value of a predetermined adjustment width may be specified.
- the information on the adjustment scheme may further include the number of adjusted audio signals to be generated.
- the adjustment method registration unit 110 may display an interface for receiving input on the display device 70 .
- Information related to typical adjustment methods may be stored in the adjustment method storage unit 180 in advance. In such a case, the function of the adjustment method registration unit 110 is unnecessary except when changing the adjustment method.
- the adjusted waveform group generation unit 120 generates a plurality of adjusted audio signals by adjusting the audio signal input from the audio signal acquisition unit 100 according to the information related to the adjustment method stored in the adjustment method storage unit 180.
- the adjustment method storage unit 180 stores information related to an adjustment method in which the adjustment width is ⁇ 1.0% with the speech rate as an attribute to be adjusted.
- the adjusted waveform group generation unit 120 generates an adjusted speech signal with the speech rate of the input speech signal adjusted to 101%, an adjusted speech signal with 99%, an adjusted speech signal with 102%, and 98%. , respectively. If it is designated to use random numbers as the adjustment range, the adjustment range may be determined by successively obtaining the adjustment range using random numbers.
- Pitch, formant, etc. can be changed by known pitch shift and formant shift techniques such as SOLA (Synchronized Overlap-Add method) and PV (Phase Vocoder).
- SOLA Synchronized Overlap-Add method
- PV Phase Vocoder
- the SN ratio can be changed by regarding a component having a predetermined amplitude or less in the audio signal as noise and changing the magnitude of that component. Attributes of other audio signals can also be changed by known techniques.
- the number of adjusted audio signals to be generated is included in the information related to the adjustment method, the specified number of adjusted audio signals are generated. If not, a predetermined predetermined number of adjusted audio signals may be generated.
- the adjusted waveform group generation unit 120 outputs the original speech signal and the plurality of adjusted speech signals to the speech recognition unit 130 as data related to the adjusted waveform group.
- the speech recognition unit 130 performs known speech recognition on each speech signal (original speech signal and a plurality of adjusted speech signals) included in the data related to the adjusted waveform group input by the adjusted waveform group generation unit 120. process. Then, the speech recognition result for each speech signal is output to the tally result generation unit 150 .
- the speech recognition processing performed by the speech recognition unit 130 includes, for example, DP (Dynamic Programming) matching, HMM (Hidden Markov Model), GMM (Gaussian Mixture Model)-HMM, DNN (Deep Neural Network)-HMM, Known models such as RNN (Recurrent Neural Network) and LSTM (Long Short-Term Memory) may be used.
- the aggregation method registration unit 140 is associated with an aggregation method indicating by what kind of statistical processing the results of speech recognition performed on each audio signal included in the data related to the adjusted waveform group by the speech recognition unit 130 are aggregated.
- Information is received and registered in the tabulation method storage unit 190 .
- Information related to the aggregation method includes at least information related to statistical processing capable of aggregating one result based on a plurality of data.
- the information related to the counting method may be information specifying the most frequent transcribed character string in the group of transcribed character strings as a result of speech recognition.
- it may be information specifying a transcribed character string close to the median reliability of each transcribed character string as a result of speech recognition.
- the tabulation method registration unit 140 may display an interface for accepting input on the display device 70.
- information related to a typical counting method may be stored in the counting method storage unit 190 in advance. In such a case, the function of the counting method registration unit 140 is unnecessary except when changing the counting method.
- the aggregation result generation unit 150 executes predetermined statistical processing on the result of voice recognition of the data related to the adjusted waveform group by the voice recognition unit 130 according to the information related to the aggregation method stored in the aggregation method storage unit 190 . Then, the result of the statistical processing is output as the aggregate result.
- FIG. 5 shows an example in which the transcribed character string corresponding to the mode of the transcribed character string group as the result of speech recognition is specified as the information related to the aggregation method.
- the adjustment waveform group generation unit 120 When the audio signal output by the audio signal acquisition unit 100 is input to the adjusted waveform group generation unit 120, the adjustment waveform group generation unit 120 generates the input audio according to the information related to the adjustment method stored in the adjustment method storage unit 180.
- a plurality of audio signals are generated that are adjusted for predetermined attributes of the signals.
- a plurality of adjusted speech signals are generated by adjusting the speech rate by a predetermined adjustment width. Then, these audio signals and the plurality of adjusted audio signals are output to the speech recognition unit 130 as data related to the adjusted waveform group.
- the speech recognition unit 130 performs speech recognition processing on each speech signal included in the adjusted waveform group.
- the result is the transcript recognized from each audio signal and its confidence level.
- the tally result generation unit 150 performs tally processing for obtaining the transcribed character string corresponding to the mode of the transcribed character string for these speech recognition results. Since the mode value of the transcribed character string is "equipment setting", the tally result generation unit 150 outputs the transcribed character string "device setting" as the result of the tallying process.
- FIG. 6 shows an example in which a transcribed character string close to the median reliability of each transcribed character string as a result of speech recognition is specified as information related to the aggregation method.
- the adjustment waveform group generation unit 120 When the audio signal output by the audio signal acquisition unit 100 is input to the adjusted waveform group generation unit 120, the adjustment waveform group generation unit 120 generates the input audio according to the information related to the adjustment method stored in the adjustment method storage unit 180.
- a plurality of audio signals are generated that are adjusted for predetermined attributes of the signals.
- a plurality of adjusted audio signals are generated by adjusting the amplitude value of the audio signal by a predetermined adjustment width. Then, these audio signals and the plurality of adjusted audio signals are output to the speech recognition unit 130 as data related to the adjusted waveform group.
- the speech recognition unit 130 performs speech recognition processing on each speech signal included in the adjusted waveform group.
- the result is the transcript recognized from each audio signal and its confidence level.
- the tally result generation unit 150 performs tally processing for obtaining the median reliability of these speech recognition results. Assume that the median reliability is 0.81.
- the aggregation result generation unit 150 generates a transcription character string of the speech recognition result of the adjusted speech signal 4, which is the speech recognition result whose reliability value is closest to 0.81 as a result of aggregation processing. I want to reduce driving time" is output.
- the command processing unit 160 accepts the tally result output from the tally result generating unit 150 as a voice command. Then, according to the received voice command, a predetermined function corresponding to the voice command is executed.
- the predetermined function may be a general function that the control device has. For example, a function of calling a predetermined screen of the voice recognition device, a function of setting a predetermined parameter, a function of controlling the industrial machine 2, and the like are exemplified.
- the output unit 170 displays and outputs the totalization result output from the totalization result generation unit 150 on the display device 70 .
- the output unit 170 may display the aggregated result at a position (for example, a state display area at the bottom of the screen) that does not interfere with the display of a predetermined function being executed on the screen of the display device 70. .
- the information may be displayed and output in the form of a dialog or the like.
- the output unit 170 may transmit and output the aggregated result to other industrial machines 4 , the fog computer 6 , the cloud server 7 , or other higher-level computers via the network 5 .
- the log may be output to a log recording area provided in advance on the nonvolatile memory 14 or the like.
- the speech recognition device 1 having the above configuration generates a plurality of adjusted speech signals with similar waveforms for the acquired speech signal. Next, speech recognition processing is performed on the generated adjusted waveform group.
- predetermined statistical processing is performed on the result of voice recognition processing, even if predetermined attributes of the voice signal are disturbed due to environmental factors at the manufacturing site, the processing accuracy of voice recognition is robust. . Therefore, it is expected that the accuracy rate of voice recognition will also be improved.
- the present invention is not limited to the above-described examples of the embodiments, and can be implemented in various modes by adding appropriate modifications.
- the speech recognition device 1 has all the functions.
- some functions may be provided on other computers such as the fog computer 6 and the cloud server 7 .
- an adjustment method registration unit 110, an aggregation method registration unit 140, an adjustment method storage unit 180, and an aggregation method storage unit 190 are provided on the fog computer, and information related to the adjustment method and information related to the aggregation method are provided.
- the information may be shared and used by a plurality of speech recognition devices 1 (control devices).
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Acoustics & Sound (AREA)
- Computational Linguistics (AREA)
- Health & Medical Sciences (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- Quality & Reliability (AREA)
- Signal Processing (AREA)
- Management, Administration, Business Operations System, And Electronic Commerce (AREA)
- Programmable Controllers (AREA)
- Arrangements For Transmission Of Measured Signals (AREA)
- Measurement Of Mechanical Vibrations Or Ultrasonic Waves (AREA)
- Numerical Control (AREA)
Abstract
Description
そこで、製造現場では認識結果の乱れに対応できる音声認識の技術が望まれている。
図1は本発明の一実施形態による音声認識装置の要部を示す概略的なハードウェア構成図である。本実施形態による音声認識装置1は、工場などの製造現場に設置された産業機械2を制御する制御装置の上に実装することができる。また、音声認識装置1は、制御装置に併設されたパソコンや、制御装置と有線乃至無線のネットワークを介して接続されたフォグコンピュータ6、クラウドサーバ7などのコンピュータ上に実装することができる。以下では、本実施形態による音声認識装置1を、産業機械2を制御する制御装置上に実装した例に基づいて説明する。
例えば、上記した実施形態では、音声認識装置1上に全ての機能を持たせている例を示している。しかしながら、一部の機能をフォグコンピュータ6やクラウドサーバ7などの他のコンピュータ上に設けるように構成してもよい。例えば、図7に例示するように、調整方式登録部110、集計方式登録部140、調整方式記憶部180、集計方式記憶部190をフォグコンピュータ上に設け、調整法式に係る情報や集計方式に係る情報を複数の音声認識装置1(制御装置)で共有して利用するようにしてもよい。
2 産業機械
4 産業機械
5 ネットワーク
6 フォグコンピュータ
7 クラウドサーバ
11 CPU
12 ROM
13 RAM
14 不揮発性メモリ
15,17,18,20,21 インタフェース
16 PLC
19 I/Oユニット
22 バス
30 軸制御回路
40 サーボアンプ
50 サーボモータ
70 表示装置
71 入力装置
72 外部機器
73 音声センサ
100 音声信号取得部
110 調整方式登録部
120 調整波形群生成部
130 音声認識部
140 集計方式登録部
150 集計結果生成部
160 コマンド処理部
170 出力部
180 調整方式記憶部
190 集計方式記憶部
Claims (9)
- 製造現場において入力された音声信号を音声認識して音声コマンドとして利用する音声認識装置であって、
入力された音声信号の所定の属性に対して複数の異なる調整を行い、これに対応する複数の調整済み音声信号を生成する調整波形群生成部と、
前記調整波形群生成部が出力する前記音声信号及び複数の前記調整済み音声信号に対する音声認識を行う音声認識部と、
を備え、
前記調整波形群生成部が行う調整は、調整対象の属性として発話速度を含む、
音声認識装置。 - 前記調整波形群生成部が行う調整は、前記調整対象の属性に対して乱数によって決まる変更を加えるものである、
請求項1に記載の音声認識装置。 - 前記音声信号及び複数の前記調整済み音声信号に対して、前記音声認識部が認識した認識結果群を所定の集計方式で統計処理する集計結果生成部を更に備える、
請求項1または2に記載の音声認識装置。 - 前記集計結果生成部は、書き起こし結果文字列群の最頻値を出力する、
請求項3に記載の音声認識装置。 - 前記集計結果生成部は、書き起こし結果信頼度群の中央値を出力する、
請求項3に記載の音声認識装置。 - 前記集計結果生成部が統計処理した結果をユーザに提示する出力部を更に備える、
請求項3~5のいずれか1つに記載の音声認識装置。 - 調整対象となる前記属性とその調整幅について、ユーザ入力を受け付け登録する調整方式登録部をさらに備える、
請求項1~6のいずれか1つに記載の音声認識装置。 - 前記集計方式について、ユーザ入力を受け付け登録する集計方式登録部をさらに備える、
請求項3~6のいずれか1つに記載の音声認識装置。 - 製造現場において入力された音声信号を音声認識して音声コマンドとして利用する音声認識装置で実行されるプログラムを記録するコンピュータ読み取り可能な記録媒体であって、
入力された音声信号の発話速度を含む所定の属性に対して複数の異なる調整を行い、これに対応する複数の調整済み音声信号を生成する調整波形群生成部、
前記調整波形群生成部が出力する前記音声信号及び複数の前記調整済み音声信号に対する音声認識を行う音声認識部、
としてコンピュータを機能させるプログラムを記録するコンピュータ読み取り可能な記録媒体。
Priority Applications (5)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202280090625.4A CN118648057A (zh) | 2022-02-08 | 2022-02-08 | 声音识别装置以及计算机可读取的记录介质 |
| PCT/JP2022/004938 WO2023152803A1 (ja) | 2022-02-08 | 2022-02-08 | 音声認識装置、及びコンピュータが読み取り可能な記録媒体 |
| US18/833,395 US20250124917A1 (en) | 2022-02-08 | 2022-02-08 | Voice recognition device and computer-readable recording medium |
| DE112022005608.8T DE112022005608T5 (de) | 2022-02-08 | 2022-02-08 | Spracherkennungsvorrichtung und computerlesbares aufzeichnungsmedium |
| JP2023579897A JP7712405B2 (ja) | 2022-02-08 | 2022-02-08 | 音声認識装置、及びコンピュータが読み取り可能な記録媒体 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2022/004938 WO2023152803A1 (ja) | 2022-02-08 | 2022-02-08 | 音声認識装置、及びコンピュータが読み取り可能な記録媒体 |
Publications (2)
| Publication Number | Publication Date |
|---|---|
| WO2023152803A1 true WO2023152803A1 (ja) | 2023-08-17 |
| WO2023152803A9 WO2023152803A9 (ja) | 2024-06-06 |
Family
ID=87563808
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2022/004938 Ceased WO2023152803A1 (ja) | 2022-02-08 | 2022-02-08 | 音声認識装置、及びコンピュータが読み取り可能な記録媒体 |
Country Status (5)
| Country | Link |
|---|---|
| US (1) | US20250124917A1 (ja) |
| JP (1) | JP7712405B2 (ja) |
| CN (1) | CN118648057A (ja) |
| DE (1) | DE112022005608T5 (ja) |
| WO (1) | WO2023152803A1 (ja) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH09325798A (ja) * | 1996-06-06 | 1997-12-16 | Matsushita Electric Ind Co Ltd | 音声認識装置 |
| JP2004347956A (ja) * | 2003-05-23 | 2004-12-09 | Toshiba Corp | 音声認識装置、音声認識方法及び音声認識プログラム |
| JP2006330389A (ja) * | 2005-05-26 | 2006-12-07 | Matsushita Electric Works Ltd | 音声認識装置 |
| JP2015215503A (ja) * | 2014-05-12 | 2015-12-03 | 日本電信電話株式会社 | 音声認識方法、音声認識装置および音声認識プログラム |
Family Cites Families (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5893062A (en) * | 1996-12-05 | 1999-04-06 | Interval Research Corporation | Variable rate video playback with synchronized audio |
| US6336092B1 (en) * | 1997-04-28 | 2002-01-01 | Ivl Technologies Ltd | Targeted vocal transformation |
| US6598020B1 (en) * | 1999-09-10 | 2003-07-22 | International Business Machines Corporation | Adaptive emotion and initiative generator for conversational systems |
| US10403269B2 (en) * | 2015-03-27 | 2019-09-03 | Google Llc | Processing audio waveforms |
-
2022
- 2022-02-08 JP JP2023579897A patent/JP7712405B2/ja active Active
- 2022-02-08 US US18/833,395 patent/US20250124917A1/en active Pending
- 2022-02-08 WO PCT/JP2022/004938 patent/WO2023152803A1/ja not_active Ceased
- 2022-02-08 CN CN202280090625.4A patent/CN118648057A/zh active Pending
- 2022-02-08 DE DE112022005608.8T patent/DE112022005608T5/de active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH09325798A (ja) * | 1996-06-06 | 1997-12-16 | Matsushita Electric Ind Co Ltd | 音声認識装置 |
| JP2004347956A (ja) * | 2003-05-23 | 2004-12-09 | Toshiba Corp | 音声認識装置、音声認識方法及び音声認識プログラム |
| JP2006330389A (ja) * | 2005-05-26 | 2006-12-07 | Matsushita Electric Works Ltd | 音声認識装置 |
| JP2015215503A (ja) * | 2014-05-12 | 2015-12-03 | 日本電信電話株式会社 | 音声認識方法、音声認識装置および音声認識プログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| JPWO2023152803A1 (ja) | 2023-08-17 |
| US20250124917A1 (en) | 2025-04-17 |
| DE112022005608T5 (de) | 2024-10-02 |
| JP7712405B2 (ja) | 2025-07-23 |
| WO2023152803A9 (ja) | 2024-06-06 |
| CN118648057A (zh) | 2024-09-13 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP4657736B2 (ja) | ユーザ訂正を用いた自動音声認識学習のためのシステムおよび方法 | |
| US20150242182A1 (en) | Voice augmentation for industrial operator consoles | |
| TWI684177B (zh) | 工作機械語音控制系統 | |
| US11474497B2 (en) | Numerical control device, machine learning device, and numerical control method | |
| EP3232380A1 (en) | Device maintenance apparatus, method for maintaining device, and storage medium | |
| CN118351847A (zh) | 一种语音交互式机床控制系统及其控制方法 | |
| JP2020021124A (ja) | 制御装置 | |
| JP7712405B2 (ja) | 音声認識装置、及びコンピュータが読み取り可能な記録媒体 | |
| US20240282310A1 (en) | Speech recognition device | |
| KR102344426B1 (ko) | 표면 실장 부품 조립 장비의 작동 오류 검출 장치 | |
| CN113383281B (zh) | 控制装置以及存储介质 | |
| JPH03120598A (ja) | 音声認識方法及び装置 | |
| JP7199616B1 (ja) | 加工結果評価装置、加工結果評価方法、加工条件決定装置、および加工条件決定方法 | |
| JP7820405B2 (ja) | 音声認識装置、およびコンピュータ読み取り可能な記憶媒体 | |
| JP7849466B2 (ja) | 機械操作装置 | |
| WO2023042277A1 (ja) | 操作訓練装置、操作訓練方法、およびコンピュータ読み取り可能な記憶媒体 | |
| Baalman | The Machine is learning | |
| JP6452826B2 (ja) | ファクトリーオートメーションシステムおよびリモートサーバ | |
| KR100802483B1 (ko) | 모션 및 피엘씨 통합제어장치 | |
| JP7576729B1 (ja) | 制御システム、制御方法及びプログラム | |
| TWI843084B (zh) | 控制系統、資訊處理方法以及資訊處理裝置 | |
| Pai et al. | Implementation of a Voice-ControlSystem for Issuing Commands in a Virtual Manufacturing Simulation Process | |
| KR20260044366A (ko) | 명령어 인식 시스템 및 방법 | |
| WO2023139769A1 (ja) | 文法調整装置、及びコンピュータが読み取り可能な記憶媒体 | |
| WO2023139770A1 (ja) | 文法作成支援装置、及びコンピュータが読み取り可能な記憶媒体 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 22925825 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2023579897 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 18833395 Country of ref document: US |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 202280090625.4 Country of ref document: CN |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 22925825 Country of ref document: EP Kind code of ref document: A1 |
|
| WWP | Wipo information: published in national office |
Ref document number: 18833395 Country of ref document: US |