WO2007108355A1 - 情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体 - Google Patents

情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体 Download PDF

Info

Publication number
WO2007108355A1
WO2007108355A1 PCT/JP2007/054881 JP2007054881W WO2007108355A1 WO 2007108355 A1 WO2007108355 A1 WO 2007108355A1 JP 2007054881 W JP2007054881 W JP 2007054881W WO 2007108355 A1 WO2007108355 A1 WO 2007108355A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
sound source
source information
input
information processing
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2007/054881
Other languages
English (en)
French (fr)
Inventor
Kiyoshi Morikawa
Koji Koga
Hiroaki Shibasaki
Mari Kitada
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Pioneer Corp
Original Assignee
Pioneer Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Pioneer Corp filed Critical Pioneer Corp
Publication of WO2007108355A1 publication Critical patent/WO2007108355A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G09EDUCATION; CRYPTOGRAPHY; DISPLAY; ADVERTISING; SEALS
    • G09BEDUCATIONAL OR DEMONSTRATION APPLIANCES; APPLIANCES FOR TEACHING, OR COMMUNICATING WITH, THE BLIND, DEAF OR MUTE; MODELS; PLANETARIA; GLOBES; MAPS; DIAGRAMS
    • G09B5/00Electrically-operated educational appliances
    • G09B5/04Electrically-operated educational appliances with audible presentation of the material to be studied
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L25/00Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
    • G10L25/48Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use
    • G10L25/69Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00 specially adapted for particular use for evaluating synthetic or decoded voice signals

Definitions

  • Information processing apparatus information processing method, information processing program, and computer-readable recording medium
  • the present invention relates to an information processing apparatus, an information processing method, an information processing program, and a computer-readable recording medium that process audio information.
  • the use of the present invention is not limited to the above-described information processing apparatus, information processing method, information processing program, and computer-readable recording medium.
  • a language learning apparatus using speech recognition technology outputs sound source information about English vocabulary or sentences recorded on a recording medium with standard pronunciation from a speaker. Then, the language learning device allows the user to input a foreign language vocabulary or sentence corresponding to the output sound source information from the microphone. Then, the language learning device compares the speech information input by the user with the sound source information using speech recognition technology, and calculates a score based on the matching degree between the speech information and the sound source information (for example, (See Patent Document 1).
  • Patent Document 1 Japanese Patent Application Laid-Open No. 2003-228279
  • the language learning device does not display the score on the display screen when the degree of coincidence between the speech information and the sound source information is lower than a predetermined threshold value, and again uses the same vocabulary or sentence
  • the user is made to input voice information. If the degree of coincidence between the voice information and the sound source information is equal to or greater than a predetermined threshold, the user is caused to pronounce the same vocabulary or sentence until the number of times equal to or greater than the predetermined score reaches the predetermined number. Therefore, some users may practice pronunciation with the same vocabulary or text, and the problem is that they cannot practice pronunciation for any vocabulary or text.
  • the apparatus is an information processing apparatus that is mounted on a moving body and processes audio information input from sound collection means, and reproduces sound source information recorded on a recording medium, and is reproduced by the reproduction means.
  • Output means for outputting the sound source information, and input means for allowing the user to input the sound information corresponding to the sound source information from the sound collection means after the sound source information is output by the output means.
  • calculating means for calculating a difference amount between the audio information input by the input means and the sound source information, and evaluating whether the difference amount calculated by the calculating means is equal to or greater than a predetermined amount. Evaluation means; and control means for reproducing again the sound source information reproduced by the reproduction means when the difference amount is equal to or larger than the predetermined amount as a result of the evaluation by the evaluation means. And butterflies.
  • the information processing method according to the invention of claim 6 is an information processing method that is mounted on a mobile body and processes audio information input from sound collection means, and is recorded on a recording medium.
  • a reproduction step of reproducing information; an output step of outputting the sound source information reproduced by the reproduction step; and the sound information corresponding to the sound source information is used after the sound source information is output by the output step.
  • An input process that the user inputs from the sound collecting means, a calculation step for calculating a difference amount between the sound information and the sound source information input in the input step, and the difference amount calculated in the calculation step If the difference amount is equal to or greater than the predetermined amount as a result of the evaluation step for evaluating whether or not the power is equal to or greater than a predetermined amount, and the evaluation by the evaluation step, the sound source information reproduced by the reproduction step is Characterized by comprising a control step of degrees play, the.
  • an information processing program according to claim 7 causes a computer to execute the information processing method according to claim 6.
  • a computer-readable recording medium according to the invention of claim 8 is characterized in that the information processing program according to claim 7 is recorded.
  • FIG. 1 is a block diagram showing an example of a functional configuration of an information processing apparatus according to an embodiment.
  • FIG. 2 is a flowchart showing the contents of the processing of the information processing apparatus that is effective in the present embodiment. It is
  • FIG. 3 is a block diagram showing an example of a hardware configuration of a navigation device that works well with the present embodiment.
  • FIG. 4 is a flowchart showing the contents of processing of the navigation device according to the present embodiment.
  • FIG. 1 is a block diagram showing an example of a functional configuration of the information processing apparatus according to the present embodiment.
  • the information processing apparatus 100 includes a reproduction unit 101, an output unit 102, an input unit 103, a calculation unit 104, an evaluation unit 105, and a control unit 106.
  • the playback unit 101 plays back the sound source information recorded on the recording medium.
  • the sound source information is information related to a vocabulary or sentence in a foreign language and is recorded on a recording medium with a standard pronunciation.
  • the output unit 102 outputs the sound source information being played back by the playback unit 101.
  • Output unit 102 Is a speaker or the like, and outputs information related to a vocabulary or a sentence of a foreign language being reproduced by the reproduction unit 101.
  • the input unit 103 causes the user to input sound information corresponding to the sound source information from the sound collecting means.
  • the sound collecting means is, for example, a microphone, and the input unit 103 allows the user to input the output foreign language vocabulary or text by voice from the microphone after the foreign language vocabulary or text is output from the speaker.
  • the calculation unit 104 calculates a difference amount between the audio information input by the input unit 103 and the sound source information. At this time, the calculation unit 104 calculates the difference from the sound source information by excluding the audio information other than the audio information emitted by the user from the audio information input by the input unit 103. Specifically, the calculation unit 104 calculates the amount of difference by analyzing the degree of coincidence between the voice information input from the microphone by the user and the sound source information using voice recognition technology. At this time, the calculation unit 104 calculates the difference amount by excluding voice information related to noise generated when the moving body travels, sound generated when the air conditioner of the moving body operates, and the like from the voice information input from the microphone.
  • the evaluation unit 105 evaluates whether or not the difference amount calculated by the calculation unit 104 is greater than or equal to a predetermined amount.
  • the predetermined amount of difference is preliminarily stored in a storage unit (not shown) and may be set by the user.
  • the control unit 106 causes the sound source information reproduced by the reproduction unit 101 to be reproduced again when the difference amount calculated by the calculation unit 104 is greater than or equal to a predetermined amount.
  • the control unit 106 causes the reproduction unit 101 to reproduce other sound source information recorded on the recording medium.
  • the control unit 106 causes the reproduction unit 101 to reproduce other sound source information recorded on the recording medium.
  • FIG. 2 is a flowchart showing the contents of processing of the information processing apparatus according to this embodiment.
  • the information processing apparatus 100 determines whether or not an operation start request has been received (step S201).
  • the operation start request can be made, for example, when the user inputs an instruction by operating an operation unit (not shown). Kona.
  • step S201 the reproduction unit 101 waits until an operation start request is accepted (step S201: No loop). If accepted (step S201: Yes), the playback unit 101 stores the sound source information recorded on the recording medium. Play (step S202).
  • the sound source information is information related to vocabulary or sentences in a foreign language, and is recorded on a recording medium with standard pronunciation.
  • the output unit 102 outputs the sound source information reproduced by the reproduction unit 101 in step S202 (step S203).
  • the output unit 102 is a speaker or the like, and outputs information related to a vocabulary or a sentence of a foreign language being reproduced by the reproduction unit 101.
  • the input unit 103 causes the user to input sound information corresponding to the output sound source information from the sound collecting means (step S204).
  • the sound collecting means is, for example, a microphone, and the input unit 103 outputs the foreign language vocabulary or text to the user from the microphone after the foreign language vocabulary or text is output. To input.
  • the calculation unit 104 calculates a difference amount between the audio information input by the input unit 103 and the sound source information (step S205). At this time, the calculation unit 104 calculates the difference from the sound source information by excluding the audio information other than the audio information emitted by the user from the audio information input by the input unit 103. Specifically, the calculation unit 104 calculates the amount of difference by analyzing the degree of coincidence between the voice information input from the microphone by the user and the sound source information using voice recognition technology. At this time, the calculation unit 104 calculates the difference amount by excluding voice information related to noise generated when the moving body travels, sound generated when the air conditioner of the moving body operates, and the like from the voice information input from the microphone.
  • the evaluation unit 105 evaluates whether or not the difference amount calculated by the calculation unit 104 in step S205 is a predetermined amount or more (step S206).
  • the predetermined amount of difference is stored in a storage unit (not shown) and may be set by the user. If the difference amount is not greater than or equal to the predetermined amount in step S206 (step S206: No), the information processing apparatus 100 proceeds to the process of step S208 and determines whether or not an operation end request has been received (step S208). ).
  • step S206: Yes it is determined whether or not the same sound source information has been reproduced a predetermined number of times.
  • the predetermined number of times is stored in advance in a storage unit (not shown) and may be set by the user. If the same sound source information has not been reproduced a predetermined number of times in step S207 (step S207: No), the information processing apparatus 100 returns to the process of step S202 and repeats the process.
  • step S207: Yes information processing apparatus 100 determines whether or not an operation end request has been received (step S208).
  • the operation end request is made when the user inputs an instruction by operating an operation unit (not shown). If an operation end request is accepted in step S208 (step S208: Yes), the information processing apparatus 100 ends a series of processing.
  • step S208 determines whether the operation end request is received in step S208 (step S208: No). If the operation end request is not received in step S208 (step S208: No), control unit 106 reproduces another sound source information recorded on the recording medium (step S209). ), Return to the process of step S203, and repeat the process.
  • the information processing apparatus After reproducing the sound source information recorded on the recording medium, the information processing apparatus provides the user with audio information corresponding to the reproduced sound source information. To input. Then, the information processing apparatus calculates the difference amount between the sound source information and the sound information, and when the difference amount is equal to or larger than the predetermined amount, reproduces the same sound source information again. Further, when the difference amount is not equal to or greater than the predetermined amount, or when the same sound source information is reproduced a predetermined number of times, the information processing apparatus reproduces another sound source information recorded on the recording medium.
  • FIG. 3 is a block diagram illustrating an example of a hardware configuration of the navigation device according to the present embodiment.
  • the navigation apparatus 300 is mounted on a moving body such as a vehicle, and includes a CPU 301, a ROM 302, a RAM 303, a magnetic disk drive 304, a magnetic disk 305, an optical disk drive 306, and the like. , Optical disk 307, audio I / F (interface) 308, microphone 309, speaker 310, human power device 311, video I / F 312, display 3 13, communication I / F 314, GPS unit 315 And various sensors 316. Each component 30 :! to 316 is connected by a bus 320.
  • the CPU 301 controls the entire navigation device 300.
  • the ROM 302 records programs such as a boot program, a current position calculation program, a route search program, a route guidance program, a voice generation program, a voice recognition program, a map information display program, a communication program, a database creation program, and a data analysis program. Les.
  • the RAM 303 is used as a work area for the CPU 301.
  • the current position calculation program calculates the current position of the vehicle (the current position of the navigation device 300) based on output information from a GPS unit 315 and various sensors 316 described later.
  • the route search program searches for an optimal route from the departure point to the destination point using map information or the like recorded on the optical disk 307 to be described later.
  • the optimal route is the shortest (or fastest) route to the destination or the route that best meets the conditions specified by the user.
  • not only the destination point but also a route to a stop point or a resting point may be searched.
  • the guidance route searched by executing the route search program is output to the audio IZF 308 and the video I / F 312 via the CPU 301.
  • the route guidance program is read from the guidance route information searched by executing the route search program, the current position information of the vehicle calculated by executing the current position calculation program, and the optical disc 307. Real-time route guidance information is generated based on map information. By executing the route guidance program The route guidance information generated in this way is output to the audio I / F 308 and the video I / F 312 via the CPU 301.
  • the sound generation program generates tone and sound information corresponding to the pattern. That is, based on the route guidance information generated by executing the route guidance program, the virtual sound source corresponding to the guidance point is set and the voice guidance information is generated, and the voice I / F 308 is sent via the CPU 301. Output.
  • the voice recognition program recognizes voice information input by the user from a microphone 309 described later.
  • the speech recognition program performs speech recognition by removing information related to noise other than the user's speech from the speech information input from the microphone 309.
  • Information relating to noise is, for example, information relating to vehicle engine sound, tire friction sound, wind noise, sound from an audio device, or sound of an air conditioner operating.
  • the voice recognition program analyzes a spectrum of voice information input from the microphone 309. At this time, the voice recognition program corrects the spectrum of the voice information based on the spectrum of the information regarding the noise input from the microphone 309 before and after the user's voice information is input from the microphone 309. Thereby, the voice recognition program can recognize the spectrum of only the user's voice information input from the microphone 309.
  • the map information display program determines the display format of the map information displayed on the display 313 by the video I / F 312 and displays the map information on the display 313 according to the determined display format.
  • the CPU 301 When the CPU 301 receives a request to start the shadowing mode, the CPU 301 sets the start address for the sound source information recorded on the magnetic disk 305 or the optical disk 307, and reproduces the sound source information from the start address.
  • the sound source information is, for example, information related to vocabulary and sentences in a foreign language recorded on the magnetic disk 305 or the optical disk 307 with standard pronunciation. Specifically, the sound source information is reproduced by causing the CPU 301 to output the sound source information to the speaker 310 via the audio IZF 308 described later. Then, when it is detected that the sound source information is silent by various sensors 316 described later, the CPU 301 temporarily ends the sound source information and sets an end address.
  • the CPU 301 After audio information is input from a microphone 309, which will be described later, by the user, the CPU 301 compares the sound source information reproduced from the start address to the end address with the audio information input from the microphone 309, and the difference amount Is calculated. When the CPU 301 detects that the difference amount is equal to or larger than the predetermined amount by the various sensors 316, the CPU 301 reproduces the sound source information from the start address to the end address again. In addition, when the various sensors 316 detect that the difference amount is equal to or less than the predetermined amount, and when the sound source information from the start address to the end address is reproduced a predetermined number of times, the CPU 301 sets the end address to the start address again. The sound source information is reproduced from the start address.
  • the magnetic disk drive 304 controls reading / writing of data with respect to the magnetic disk 305 according to the control of the CPU 301.
  • the magnetic disk 305 records data written under the control of the magnetic disk drive 304.
  • HD node disk
  • FD flexible disk
  • the optical disk drive 306 controls reading / writing of data with respect to the optical disk 307 according to the control of the CPU 301.
  • the optical disk 307 is a detachable recording medium from which data is read according to the control of the optical disk drive 306.
  • a writable recording medium can be used as the optical disk 307.
  • the removable recording medium may be the power of the optical disk 307, MO, memory card, or the like.
  • map information recorded on a recording medium such as the magnetic disk 305 or the optical disk 307 is map information used for route search / route guidance.
  • the map information includes background data that represents features (features) such as buildings, rivers, and the ground surface, and road shape data that represents the shape of the road. Two-dimensional or three-dimensional data is displayed on the display screen of the display 313. Drawn on. Examples of other information recorded on a recording medium such as the magnetic disk 305 and the optical disk 307 include sound source information related to foreign vocabulary and sentences recorded with standard pronunciation.
  • the road shape data further includes traffic condition data.
  • the traffic condition data includes, for example, the presence or absence of traffic lights and pedestrian crossings for each node, the presence or absence of entrances and junctions on highways, the length (distance) for each link, road width, direction of travel, road type (high speed Road, toll road, general road, etc.).
  • the traffic condition data stores past congestion information obtained by statistically processing the past congestion information based on the season * day of the week and large consecutive holidays' time.
  • the navigation device 300 obtains information on the current traffic jam from the road traffic information received by the communication I / F 314 described later. It becomes.
  • map information, sound source information, and the like are recorded on a recording medium such as the magnetic disk 305 and the optical disk 307.
  • the present invention is not limited to this.
  • the map information and the sound source information may be provided outside the navigation device 300, which is not integrated with the hardware of the navigation device 300 and recorded only. Good.
  • the navigation device 300 acquires map information and sound source information via the network through the communication IZF 314, for example.
  • the acquired map information and sound source information are stored in the RAM 303 or the like.
  • Audio I / F 308 is connected to microphone 309 for audio input and speaker 310 for audio output.
  • the sound received by the microphone 309 is A / D converted in the sound I / F 308.
  • sound is output from the speaker 310.
  • the sound input from the microphone 309 can be recorded on the magnetic disk 305 or the optical disk 307 as sound data.
  • the voice information input from the microphone 309 is A / D converted in the voice I / F 308 and then subjected to spectrum analysis according to the voice recognition program.
  • Examples of the input device 311 include a remote controller, a keyboard, a mouse, and a touch panel, each having a plurality of keys for inputting characters, numerical values, various instructions, and the like.
  • the video I / F 312 is connected to the display 313.
  • the video I / F 312 includes, for example, a graphic controller that controls the entire display 313, a buffer memory such as VRAM (Video RAM) that temporarily records image information that can be displayed immediately, and a graphic controller. It is composed of a control IC that controls display of the display 313 based on the output image data.
  • VRAM Video RAM
  • Display 313 displays icons, cursors, menus, windows, or various data such as characters and images.
  • this display 313 for example, a CRT, a TFT liquid crystal display, a plasma display, or the like can be adopted.
  • Display A plurality of vehicles 313 may be provided, for example, for a driver and for a passenger seated in a rear seat.
  • the communication I / F 314 is connected to a network via radio and functions as an interface between the navigation device 300 and the CPU 301.
  • the communication IZF 314 is connected to a communication network such as the Internet via radio, and also functions as an interface between the communication network and the CPU 301.
  • Communication networks include LANs, WANs, public line networks, mobile phone networks, and the like.
  • the communication I / F 314 is composed of, for example, an FM tuner, a VICS (Vehicle Information and Communication System) / beacon receiver, a wireless navigation device, and other navigation devices, and is congested delivered from the VICS center.
  • Road traffic information such as traffic regulations.
  • VICS is a registered trademark.
  • the 0-3 unit 315 receives radio waves from GPS satellites and calculates information indicating the current position of the vehicle.
  • the output information of the GPS unit 315 is used when the CPU 301 calculates the current position of the vehicle together with output values of various sensors 316 described later.
  • the information indicating the current position is information that identifies one point on the map information, such as latitude'longitude and altitude.
  • Various sensors 316 output information such as a vehicle speed sensor, an acceleration sensor, and an angular velocity sensor that can determine the position and behavior of the vehicle.
  • the output values of the various sensors 316 are used for calculation of the current position by the CCU 301, measurement of speed and direction change, and the like.
  • the various sensors 316 detect whether or not the sound source information being reproduced by the CPU 301 has become silent. Specifically, the various sensors 316 determine whether or not the output level of the sound source information being reproduced by the CPU 301 is equal to or less than a predetermined threshold value.
  • the predetermined threshold value is set in advance in a voice recognition program recorded in the R0M302.
  • the various sensors 316 detect whether or not the difference amount between the sound source information calculated by the CPU 301 and the sound information input from the microphone 309 is greater than or equal to a predetermined amount.
  • the playback unit 101 is based on the CPU 301, the magnetic disk 305, the optical disk 307, and the audio IZF 308, and the output unit 102 is based on the speaker 310.
  • the human power unit 103 and the calculation unit 104 are controlled by the microphone 309.
  • the control unit 106 realizes the respective functions by the CPU 301 and the evaluation unit 105 by the various sensors 316.
  • FIG. 4 is a flowchart showing the contents of the processing of the navigation device that is useful in this embodiment.
  • the CPU 301 of the navigation device 300 determines whether or not a request for starting a shadowing mode has been received (step S401).
  • the request to start the shadowing mode is made, for example, when the user inputs an instruction by operating an operation unit (not shown).
  • step S401 the process waits until the shadowing mode start request is accepted (step S401: No loop).
  • a start address is set for the sound source information (step S402).
  • the sound source information is, for example, information related to vocabulary and sentences in a foreign language recorded with standard pronunciation.
  • the start address may be, for example, the first track of the sound source information, or may be the position where the user has requested the start of the shadowing mode during playback of the sound source information.
  • the CPU 301 reproduces the sound source information from the start address power set in step S402 (step S403). Specifically, the CPU 301 outputs sound source information to the speaker 310 via the audio I / F 308.
  • the CPU 301 determines whether or not the sound source information being reproduced has become silent in step S402 (step S404). Specifically, the CPU 301 uses the various sensors 316 to determine whether or not the output level of the sound source information being reproduced has become a predetermined threshold value or less.
  • the predetermined threshold is preset in the voice recognition program recorded in the ROM 302.
  • step S404 when the sound source information is not silenced (step S404: No), the CPU 301 returns to the process of step S403 and repeats the process. On the other hand, when the sound source information becomes silent in step S404 (step S404: Yes), the CPU 301 pauses the sound source information being reproduced and sets the end address (step S405). [0063] Further, the CPU 301 causes the audio I / F 308 to output guidance information from the speaker 310 (step S406).
  • the guide information is information indicating that the user inputs voice information corresponding to the sound source information reproduced from the start address to the end address from the microphone 309 by voice. At this time, the CPU 301 starts recording audio information input from the microphone 309 via the audio I / F 308. Audio information input from the microphone 309 is recorded in the RAM 303, for example.
  • CPU 301 determines whether or not audio information is input from microphone 309 by the user (step S407). Specifically, when the CPU 301 detects that the input level of the audio information input from the microphones by the various sensors 316 has changed from a state below a predetermined threshold to a state above a predetermined threshold, the CPU 301 detects the user from the microphone 309. Judge that the input of audio information has started. When the CPU 301 detects that the input level of the voice information input from the microphone by the various sensors 316 has changed from a state equal to or higher than the predetermined threshold to a level equal to or lower than the predetermined threshold, Judge that input is complete.
  • the predetermined threshold of the input level is set by the CPU 301 based on information regarding noise other than the user's voice input from the microphone 309.
  • step S407 If no audio information is input from microphone 309 in step S407 (step S407: No), the process returns to step S406 until CPU 301i, and the process is repeated.
  • step S407 If audio information is input from the microphone 309 in step S407 (step S407: Yes), the CPU 301 compares the sound source information reproduced from the start address to the end address with the audio information input from the microphone 309. Thus, the difference amount is calculated (step S408). Specifically, the CPU 301 calculates the difference between the sound source information and the sound information by performing a spatial analysis of the sound source information and the sound information according to the sound recognition program recorded in the ROM 302. . At this time, the CPU 301 calculates the difference from the sound source information by excluding information related to noise other than the user's voice in the voice information.
  • step S409 determines whether or not the difference amount calculated in step S408 by various sensors 316 is equal to or greater than a predetermined amount (step S409).
  • the predetermined amount of difference is recorded in the RAM 303, for example. If the difference amount is not greater than or equal to the predetermined amount in step S409 (step S409: No), the CPU 301 proceeds to the process of step S412. Then, the end address set in step S405 is set as the start address (step S412).
  • the CPU 301 deletes the audio information input from the microphone 309 (step S410). Specifically, the audio information input from the microphone 309 by the user of the CPU 301 and recorded in the RAM 303 is deleted (step S410).
  • step S411 determines whether or not the sound source information from the start address to the end address has been reproduced a predetermined number of times.
  • step S411 when the sound source information has not been reproduced a predetermined number of times (step S411: No), the CPU 301 returns to the process of step S403 and repeats the process.
  • step S411 determines whether or not a shadowing mode end request has been received (step S413).
  • the request to end the shadowing mode is made, for example, when the user inputs an instruction by operating an operation unit (not shown).
  • step S413 when the shadowing mode termination request is not accepted
  • Step S413: No CPU 301 i is returned to the process of Step S403, and the process is repeated.
  • step S413: Yes when a request to end the shadowing mode is received in step S413 (step S413: Yes), the CPU 301 ends a series of processes.
  • the navigation device calculates the difference between the reproduced sound source information and the voice information input from the microphone by the user in the shadowing mode. If the calculated difference amount is equal to or larger than the predetermined amount, the navigation device reproduces the same sound source information again. In addition, when the calculated difference amount is equal to or smaller than the predetermined amount, and when the same sound source information is reproduced a predetermined number of times, the navigation device reproduces another sound source information. Therefore, even users who cannot pronounce vocabulary and sentences in foreign languages can improve their ability to hear foreign languages by using the shadowing mode to pronounce all vocabulary and sentences in foreign languages.
  • the information processing method described in the present embodiment can be realized by executing a program prepared in advance on a computer such as a personal computer or a workstation.
  • This program is recorded on a computer-readable recording medium such as a hard disk, a flexible disk, a CD-ROM, M0, or a DVD, and is executed by being read from the recording medium by the computer.
  • the program may be a transmission medium that can be distributed through a network such as the Internet.

Landscapes

  • Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Signal Processing (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Human Computer Interaction (AREA)
  • Computational Linguistics (AREA)
  • Acoustics & Sound (AREA)
  • Multimedia (AREA)
  • Business, Economics & Management (AREA)
  • Educational Administration (AREA)
  • Educational Technology (AREA)
  • General Physics & Mathematics (AREA)
  • Theoretical Computer Science (AREA)
  • Navigation (AREA)

Abstract

 移動体に搭載され、集音手段から入力される音声情報を処理する情報処理装置(100)であって、再生部(101)は、記録媒体に記録されている音源情報を再生する。出力部(102)は、再生部(101)により再生されている音源情報を出力する。入力部(103)は、出力部(102)により音源情報が出力された後に、音源情報に対応する音声情報を利用者に集音手段から入力させる。算出部(104)は、入力部(103)により入力された音声情報と音源情報との差異量を算出する。評価部(105)は、算出部(104)により算出された差異量が所定量以上であるか否かを評価する。制御部(106)は、評価部(105)による評価の結果、差異量が所定量以上である場合、再生部(101)によって再生された音源情報を再度再生させる。

Description

明 細 書
情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータ に読み取り可能な記録媒体
技術分野
[0001] この発明は、音声情報を処理する情報処理装置、情報処理方法、情報処理プログ ラムおよびコンピュータに読み取り可能な記録媒体に関する。ただし、この発明の利 用は、上述した情報処理装置、情報処理方法、情報処理プログラムおよびコンビユー タに読み取り可能な記録媒体には限られない。
背景技術
[0002] 従来、音声認識技術を用いた語学学習装置は、標準的な発音で記録媒体に記録 された英語の語彙または文章に関する音源情報をスピーカから出力する。そして、語 学学習装置は、出力された音源情報に相当する外国語の語彙または文章を利用者 にマイクから音声で入力させる。そして、語学学習装置は、利用者により入力された 音声情報と音源情報とを音声認識技術を用いて比較し、音声情報と音源情報との一 致度に基づいて点数を算出する(たとえば、下記特許文献 1参照。)。
[0003] 特許文献 1 :特開 2003— 228279号公報
発明の開示
発明が解決しょうとする課題
[0004] しかしながら、上記従来技術によれば、語学学習装置は、音声情報と音源情報との 一致度が所定の閾値より低い場合は点数を表示画面上に表示せず、再度同じ語彙 または文章に対して音声情報を利用者に入力させる。また、音声情報と音源情報と の一致度が所定の閾値以上である場合は、所定の点数以上である回数が所定回数 になるまで同じ語彙または文章を利用者に発音させる。したがって、利用者によって は同じ語彙または文章ば力り発音練習することになり、あらゆる語彙または文章につ いて発音練習することができないという問題が一例として挙げられる。
課題を解決するための手段
[0005] 上述した課題を解決し、 目的を達成するため、請求項 1の発明にかかる情報処理 装置は、移動体に搭載され、集音手段から入力される音声情報を処理する情報処理 装置であって、記録媒体に記録されている音源情報を再生する再生手段と、前記再 生手段により再生されている前記音源情報を出力する出力手段と、前記出力手段に より前記音源情報が出力された後に、前記音源情報に対応する前記音声情報を利 用者に前記集音手段から入力させる入力手段と、前記入力手段により入力された前 記音声情報と前記音源情報との差異量を算出する算出手段と、前記算出手段により 算出された前記差異量が所定量以上であるか否かを評価する評価手段と、前記評 価手段による評価の結果、前記差異量が前記所定量以上である場合、前記再生手 段によって再生された前記音源情報を再度再生させる制御手段と、を備えることを特 徴とする。
[0006] また、請求項 6の発明にかかる情報処理方法は、移動体に搭載され、集音手段から 入力される音声情報を処理する情報処理方法であって、記録媒体に記録されている 音源情報を再生する再生工程と、前記再生工程により再生されている前記音源情報 を出力する出力工程と、前記出力工程により前記音源情報が出力された後に、前記 音源情報に対応する前記音声情報を利用者に前記集音手段から入力させる入力ェ 程と、前記入力工程により入力された前記音声情報と前記音源情報との差異量を算 出する算出工程と、前記算出工程により算出された前記差異量が所定量以上である か否力を評価する評価工程と、前記評価工程による評価の結果、前記差異量が前 記所定量以上である場合、前記再生工程によって再生された前記音源情報を再度 再生させる制御工程と、を含むことを特徴とする。
[0007] また、請求項 7の発明に力かる情報処理プログラムは、請求項 6に記載の情報処理 方法をコンピュータに実行させることを特徴とする。
[0008] また、請求項 8の発明にかかるコンピュータに読み取り可能な記録媒体は、請求項 7に記載の情報処理プログラムを記録したことを特徴とする。
図面の簡単な説明
[0009] [図 1]図 1は、本実施の形態にかかる情報処理装置の機能的構成の一例を示すプロ ック図である。
[図 2]図 2は、本実施の形態に力かる情報処理装置の処理の内容を示すフローチヤ ートである。
[図 3]図 3は、本実施例に力かるナビゲーシヨン装置のハードウェア構成の一例を示 すブロック図である。
[図 4]図 4は、本実施例にかかるナビゲーシヨン装置の処理の内容を示すフローチヤ ートである。
符号の説明
[0010] 100 情報処理装置
101 再生部
102 出力部
103 入力部
104 算出部
105 評価部
106 制御部
発明を実施するための最良の形態
[0011] 以下に添付図面を参照して、この発明にかかる情報処理装置、情報処理方法、情 報処理プログラムおよびコンピュータに読み取り可能な記録媒体の好適な実施の形 態を詳細に説明する。
[0012] (実施の形態)
(情報処理装置の機能的構成)
図 1を用いて、この発明にかかる情報処理装置の機能的構成について説明する。 図 1は、本実施の形態にかかる情報処理装置の機能的構成の一例を示すブロック図 である。
[0013] 図 1において、情報処理装置 100は、再生部 101と、出力部 102と、入力部 103と、 算出部 104と、評価部 105と、制御部 106と、を含み構成されている。
[0014] 再生部 101は、記録媒体に記録されている音源情報を再生する。音源情報は、外 国語の語彙または文章に関する情報であり、標準的な発音で記録媒体に記録されて いる。
[0015] 出力部 102は、再生部 101により再生されている音源情報を出力する。出力部 102 は、スピーカなどであり、再生部 101により再生されている外国語の語彙または文章 に関する情報を出力する。
[0016] 入力部 103は、出力部 102により音源情報が出力された後に、音源情報に対応す る音声情報を利用者に集音手段から入力させる。集音手段は、たとえばマイクであり 、入力部 103は、スピーカから外国語の語彙または文章が出力された後に、出力さ れた外国語の語彙または文章を利用者にマイクから音声で入力させる。
[0017] 算出部 104は、入力部 103により入力された音声情報と音源情報との差異量を算 出する。この際、算出部 104は、入力部 103により入力された音声情報のうち利用者 が発した音声情報以外の音声情報を除いて音源情報との差異量を算出する。具体 的には、算出部 104は、音声認識技術を用いて利用者によりマイクから入力された 音声情報と音源情報との一致度を解析して差異量を算出する。この際、算出部 104 は、マイクから入力された音声情報のうち移動体が走行することにより生じる騒音や 移動体の空調機が作動する音などに関する音声情報を除いて差異量を算出する。
[0018] 評価部 105は、算出部 104により算出された差異量が所定量以上であるか否かを 評価する。差異量の所定量は、図示しない記憶部にあら力じめ記憶されており、利用 者が設定してもよい。
[0019] 制御部 106は、評価部 105による評価の結果、算出部 104により算出された差異 量が所定量以上である場合、再生部 101によって再生された音源情報を再度再生さ せる。また、制御部 106は、算出部 104により算出された差異量が所定量以上でな い場合、再生部 101により記録媒体に記録されている別の音源情報を再生させる。 また、制御部 106は、同一の音源情報が所定回数再生された場合、再生部 101によ り記録媒体に記録されている別の音源情報を再生させる。
[0020] (情報処理装置の処理の内容)
つぎに、図 2を用いて、本実施の形態に力、かる情報処理装置 100の処理の内容に ついて説明する。図 2は、本実施の形態にかかる情報処理装置の処理の内容を示す フローチャートである。図 2のフローチャートにおいて、まず、情報処理装置 100は、 動作の開始要求を受け付けたか否かを判断する(ステップ S201)。動作の開始要求 は、たとえば、利用者が図示しない操作部を操作して指示入力をおこなうことによりお こなう。
[0021] ステップ S201において、動作の開始要求を受け付けるまで待って(ステップ S201 : Noのループ)、受け付けた場合 (ステップ S201 : Yes)、再生部 101は、記録媒体 に記録されている音源情報を再生する (ステップ S202)。音源情報は、外国語の語 彙または文章に関する情報であり、標準的な発音で記録媒体に記録されている。
[0022] また、出力部 102は、ステップ S202において再生部 101により再生されている音源 情報を出力する(ステップ S203)。出力部 102は、スピーカなどであり、再生部 101 により再生されている外国語の語彙または文章に関する情報を出力する。
[0023] つぎに、入力部 103は、ステップ S203において出力部 102により音源情報が出力 された後に、出力された音源情報に対応する音声情報を利用者に集音手段から入 力させる(ステップ S204)。集音手段は、たとえばマイクであり、入力部 103は、スピ 一力力、ら外国語の語彙または文章が出力された後に、出力された外国語の語彙また は文章を利用者にマイクから音声で入力させる。
[0024] そして、算出部 104は、入力部 103により入力された音声情報と音源情報との差異 量を算出する (ステップ S205)。この際、算出部 104は、入力部 103により入力され た音声情報のうち利用者が発した音声情報以外の音声情報を除いて音源情報との 差異量を算出する。具体的には、算出部 104は、音声認識技術を用いて利用者によ りマイクから入力された音声情報と音源情報との一致度を解析して差異量を算出す る。この際、算出部 104は、マイクから入力された音声情報のうち移動体が走行する ことにより生じる騒音や移動体の空調機が作動する音などに関する音声情報を除い て差異量を算出する。
[0025] つぎに、評価部 105は、ステップ S205において算出部 104により算出された差異 量が所定量以上か否かを評価する (ステップ S206)。差異量の所定量は、図示しな い記憶部にあらカ^め記憶されており、利用者が設定してもよい。ステップ S206にお いて、差異量が所定量以上でない場合 (ステップ S206 : No)、情報処理装置 100は 、ステップ S208の処理に移行し、動作の終了要求を受け付けたか否かを判断する( ステップ S208)。
[0026] 一方、ステップ S206において、差異量が所定量以上である場合 (ステップ S206 : Yes) ,同じ音源情報が所定回数再生されたか否力を判断する(ステップ S207)。所 定回数は、図示しない記憶部にあらかじめ記憶されており、利用者が設定してもよい 。ステップ S207において、同じ音源情報が所定回数再生されていない場合 (ステツ プ S207 : No)、情報処理装置 100は、ステップ S202の処理に戻り、処理を繰り返す
[0027] 一方、ステップ S207において、同じ音源情報が所定回数再生された場合 (ステツ プ S207 : Yes)、情報処理装置 100は、動作の終了要求を受け付けたか否かを判断 する(ステップ S208)。動作の終了要求は、利用者が図示しない操作部を操作して 指示入力をおこなうことによりおこなう。ステップ S208において、動作の終了要求を 受け付けた場合 (ステップ S208 : Yes)、情報処理装置 100は、一連の処理を終了す る。
[0028] 一方、ステップ S208において、動作の終了要求を受け付けていない場合 (ステツ プ S208 : No)、制御部 106は、記録媒体に記録されている別の音源情報を再生させ た後(ステップ S209)、ステップ S203の処理に戻り、処理を繰り返す。
[0029] 以上説明したように、本実施の形態によれば、情報処理装置は、記録媒体に記録 されている音源情報を再生した後に、再生した音源情報に相当する音声情報を利用 者に音声で入力させる。そして、情報処理装置は、音源情報と音声情報との差異量 を算出し、差異量が所定量以上である場合は、再度同じ音源情報を再生する。また 、差異量が所定量以上でない場合、あるいは同じ音源情報が所定回数再生された 場合は、情報処理装置は、記録媒体に記録されている別の音源情報を再生する。し たがって、同じ外国語の語彙または文章が所定回数再生された場合は、別の語彙ま たは文章が再生されるため、外国語を正しく発音できない利用者でもあらゆる語彙ま たは文章の発音練習をすることにより、外国語の聞き取り能力の向上を図ることがで きる。
実施例
[0030] 以下に、本発明の実施例について説明する。本実施例では、たとえば、車両(四輪 車、二輪車を含む)などの移動体に搭載されるナビゲーシヨン装置によって、本発明 の情報処理装置を実施した場合の一例について説明する。 [0031] (ナビゲーシヨン装置のハードウェア構成)
図 3を用いて、本実施例に力かるナビゲーシヨン装置のハードウェア構成について 説明する。図 3は、本実施例にかかるナビゲーシヨン装置のハードウェア構成の一例 を示すブロック図である。
[0032] 図 3において、ナビゲーシヨン装置 300は、車両などの移動体に搭載されており、 C PU301と、 ROM302と、 RAM303と、磁気ディスクドライブ 304と、磁気ディスク 30 5と、光ディスクドライブ 306と、光ディスク 307と、音声 I/F (インターフェース) 308と 、マイク 309と、スピーカ 310と、人力デノ イス 311と、映像 I/F312と、ディスプレイ 3 13と、通信 I/F314と、 GPSユニット 315と、各種センサ 316と、を備えてレヽる。また、 各構成部 30:!〜 316はバス 320によってそれぞれ接続されている。
[0033] まず、 CPU301は、ナビゲーシヨン装置 300全体の制御を司る。 ROM302は、ブ ートプログラム、現在位置算出プログラム、経路探索プログラム、経路誘導プログラム 、音声生成プログラム、音声認識プログラム、地図情報表示プログラム、通信プロダラ ム、データベース作成プログラム、データ解析プログラムなどのプログラムを記録して レ、る。また、 RAM303は、 CPU301のワークエリアとして使用される。
[0034] ここで、現在位置算出プログラムは、たとえば、後述する GPSユニット 315および各 種センサ 316の出力情報に基づいて、車両の現在位置(ナビゲーシヨン装置 300の 現在位置)を算出させる。
[0035] また、経路探索プログラムは、後述する光ディスク 307に記録されている地図情報 などを利用して、出発地点から目的地点までの最適な経路を探索させる。ここで、最 適な経路とは、 目的地点までの最短 (あるいは最速)経路やユーザが指定した条件 に最も合致する経路などである。また、 目的地点のみならず、立ち寄り地点や休憩地 点までの経路を探索してもよい。経路探索プログラムを実行することによって探索さ れた誘導経路は、 CPU301を介して音声 IZF308や映像 I/F312へ出力される。
[0036] また、経路誘導プログラムは、経路探索プログラムを実行することによって探索され た誘導経路情報、現在位置算出プログラムを実行することによって算出された車両 の現在位置情報、光ディスク 307から読み出された地図情報に基づいて、リアルタイ ムな経路誘導情報の生成をおこなわせる。経路誘導プログラムを実行することによつ て生成された経路誘導情報は、 CPU301を介して音声 I/F308や映像 I/F312へ 出力される。
[0037] また、音声生成プログラムは、パターンに対応したトーンと音声の情報を生成させる 。すなわち、経路誘導プログラムを実行することによって生成された経路誘導情報に 基づレ、て、案内ポイントに対応した仮想音源の設定と音声ガイダンス情報の生成を おこない、 CPU301を介して音声 I/F308へ出力する。
[0038] また、音声認識プログラムは、後述するマイク 309からユーザにより入力される音声 情報を認識する。この際、音声認識プログラムは、マイク 309から入力される音声情 報のうち、ユーザの音声以外のノイズに関する情報を除去して音声認識をおこなう。 ノイズに関する情報は、たとえば、車両のエンジン音、タイヤの摩擦音、風切り音、ォ 一ディォ装置による音、あるいは空調機が作動する音などに関する情報である。
[0039] 具体的には、音声認識プログラムは、マイク 309から入力される音声情報のスぺタト ルを分析する。この際、音声認識プログラムは、ユーザの音声情報がマイク 309から 入力される前後においてマイク 309から入力されるノイズに関する情報のスペクトル に基づいて音声情報のスペクトルを補正する。これによつて、音声認識プログラムは 、マイク 309から入力されたユーザの音声情報のみのスペクトルを認識することがで きる。
[0040] また、地図情報表示プログラムは、映像 I/F312によってディスプレイ 313に表示 する地図情報の表示形式を決定させ、決定された表示形式によって地図情報をディ スプレイ 313に表示させる。
[0041] また、 CPU301は、シャドウイングモードの開始要求を受け付けた場合、磁気ディス ク 305や光ディスク 307に記録されている音源情報に対して開始アドレスを設定し、 開始アドレスから音源情報を再生する。音源情報は、たとえば、標準の発音で磁気 ディスク 305や光ディスク 307に記録された外国語の語彙や文章に関する情報であ る。音源情報の再生は、具体的には、 CPU301が音源情報を後述する音声 IZF30 8を介してスピーカ 310に出力させることによりおこなわれる。そして、後述する各種セ ンサ 316により音源情報が無音状態なつたことを検知した場合、 CPU301は、音源 情報を一時終了して終了アドレスを設定する。 [0042] そして、ユーザにより後述するマイク 309から音声情報が入力された後、 CPU301 は、開始アドレスから終了アドレスまで再生された音源情報とマイク 309から入力され た音声情報とを比較して差異量を算出する。そして、 CPU301は、各種センサ 316 により差異量が所定量以上であることを検知した場合、 CPU301は、再度開始アドレ スから終了アドレスまで音源情報を再生する。また、各種センサ 316により差異量が 所定量以下であることを検知した場合、および開始アドレスから終了アドレスまでの 音源情報を所定回数再生した場合、 CPU301は、終了アドレスを改めて開始アドレ スに設定し、開始アドレスから音源情報を再生する。
[0043] 磁気ディスクドライブ 304は、 CPU301の制御にしたがって磁気ディスク 305に対 するデータの読み取り/書き込みを制御する。磁気ディスク 305は、磁気ディスクドラ イブ 304の制御で書き込まれたデータを記録する。磁気ディスク 305としては、たとえ ば、 HD (ノヽードディスク)や FD (フレキシブルディスク)を用レ、ること力できる。
[0044] 光ディスクドライブ 306は、 CPU301の制御にしたがって光ディスク 307に対するデ ータの読み取り/書き込みを制御する。光ディスク 307は、光ディスクドライブ 306の 制御にしたがってデータの読み出される着脱自在な記録媒体である。光ディスク 307 は、書き込み可能な記録媒体を利用することもできる。また、この着脱可能な記録媒 体として、光ディスク 307のほ力、 MO、メモリカードなどであってもよい。
[0045] 磁気ディスク 305や光ディスク 307などの記録媒体に記録される情報の一例として 、経路探索 ·経路誘導などに用いる地図情報が挙げられる。地図情報は、建物、河 川、地表面などの地物(フィーチャ)をあらわす背景データと、道路の形状をあらわす 道路形状データとを有しており、ディスプレイ 313の表示画面において 2次元または 3 次元に描画される。また、磁気ディスク 305や光ディスク 307などの記録媒体に記録 される他の情報の一例として、標準の発音で記録された外国語の語彙や文章に関す る音源情報などが挙げられる。
[0046] 道路形状データは、さらに交通条件データを有する。交通条件データには、たとえ ば、各ノードについて、信号や横断歩道などの有無、高速道路の出入り口やジャンク シヨンの有無、各リンクについての長さ(距離)、道幅、進行方向、道路種別(高速道 路、有料道路、一般道路など)などの情報が含まれている。 [0047] また、交通条件データには、過去の渋滞情報を、季節 *曜日 ·大型連休 '時刻など を基準に統計処理した過去渋滞情報を記憶している。ナビゲーシヨン装置 300は、 後述する通信 I/F314によって受信される道路交通情報によって現在発生している 渋滞の情報を得るが、過去渋滞情報により、指定した時刻における渋滞状況の予想 をおこなうことが可能となる。
[0048] なお、本実施例では地図情報や音源情報などを磁気ディスク 305や光ディスク 307 などの記録媒体に記録するようにしたが、これに限るものではなレ、。たとえば、地図情 報や音源情報は、ナビゲーシヨン装置 300のハードウェアと一体に設けられてレ、るも のに限って記録されているものではなぐナビゲーシヨン装置 300の外部に設けられ ていてもよい。その場合、ナビゲーシヨン装置 300は、たとえば、通信 IZF314を通じ て、ネットワークを介して地図情報や音源情報を取得する。取得された地図情報や音 源情報は RAM303などに記憶される。
[0049] 音声 I/F308は、音声入力用のマイク 309および音声出力用のスピーカ 310に接 続される。マイク 309に受音された音声は、音声 I/F308内で A/D変換される。ま た、スピーカ 310からは音声が出力される。なお、マイク 309から入力された音声は、 音声データとして磁気ディスク 305あるいは光ディスク 307に記録可能である。また、 マイク 309から入力された音声情報は、音声 I/F308内で A/D変換された後、音 声認識プログラムにしたがってスペクトル分析がおこなわれる。
[0050] 入力デバイス 311は、文字、数値、各種指示などの入力のための複数のキーを備 えたリモコン、キーボード、マウス、タツチパネルなどが挙げられる。
[0051] 映像 I/F312は、ディスプレイ 313と接続される。映像 I/F312は、具体的には、 たとえば、ディスプレイ 313全体の制御をおこなうグラフィックコントローラと、即時表示 可能な画像情報を一時的に記録する VRAM (Video RAM)などのバッファメモリと 、グラフィックコントローラから出力される画像データに基づいて、ディスプレイ 313を 表示制御する制御 ICなどによって構成される。
[0052] ディスプレイ 313には、アイコン、カーソル、メニュー、ウィンドウ、あるいは文字や画 像などの各種データが表示される。このディスプレイ 313は、たとえば、 CRT, TFT 液晶ディスプレイ、プラズマディスプレイなどを採用することができる。また、ディスプレ ィ 313は、車両に複数備えられていてもよぐたとえば、ドライバーに対するものと後 部座席に着座する搭乗者に対するものなどである。
[0053] 通信 I/F314は、無線を介してネットワークに接続され、ナビゲーシヨン装置 300と CPU301とのインターフェースとして機能する。また、通信 IZF314は、無線を介し てインターネットなどの通信網に接続され、この通信網と CPU301とのインターフエ一 スとしても機能する。
[0054] 通信網には、 LAN, WAN,公衆回線網や携帯電話網などがある。具体的には、 通信 I/F314は、たとえば、 FMチューナー、 VICS (Vehicle Information and Communication System) /ビーコンレシーバ、無線ナビゲーシヨン装置、および その他のナビゲーシヨン装置によって構成され、 VICSセンターから配信される渋滞 や交通規制などの道路交通情報を取得する。なお、 VICSは登録商標である。
[0055] また、 0?3ュニット315は、 GPS衛星からの電波を受信し、車両の現在位置を示す 情報を算出する。 GPSユニット 315の出力情報は、後述する各種センサ 316の出力 値とともに、 CPU301による車両の現在位置の算出に際して利用される。現在位置 を示す情報は、たとえば緯度'経度、高度などの、地図情報上の 1点を特定する情報 である。
[0056] 各種センサ 316は、車速センサや加速度センサ、角速度センサなどの、車両の位 置や挙動を判断することが可能な情報を出力する。各種センサ 316の出力値は、 CP U301による現在位置の算出や、速度や方位の変化量の測定などに用いられる。ま た、各種センサ 316は、 CPU301により再生中の音源情報が無音状態になったか否 力を検知する。具体的には、各種センサ 316は、 CPU301により再生中の音源情報 の出力レベルが所定の閾値以下の状態になったか否かを判断する。所定の閾値は 、R〇M302に記録されている音声認識プログラムなどにあらかじめ設定されている。 また、各種センサ 316は、 CPU301により算出された、音源情報とマイク 309から入 力された音声情報との差異量が所定量以上であるか否力、を検知する。
[0057] なお、実施の形態に力かる情報処理装置 100の機能的構成のうち、再生部 101は 、 CPU301や磁気ディスク 305や光ディスク 307や音声 IZF308によって、出力部 1 02は、スピーカ 310によって、人力部 103は、マイク 309によって、算出部 104と制 御部 106は、 CPU301によって、評価部 105は、各種センサ 316によって、それぞれ の機能を実現する。
[0058] (ナビゲーシヨン装置 300の処理の内容)
つぎに、図 4を用いて、本実施例にかかるナビゲーシヨン装置 300の処理の内容に ついて説明する。図 4は、本実施例に力かるナビゲーシヨン装置の処理の内容を示 すフローチャートである。図 4のフローチャートにおいて、まず、ナビゲーシヨン装置 3 00の CPU301は、シャドウイングモードの開始要求を受け付けたか否かを判断する( ステップ S401)。シャドウイングモードの開始要求は、たとえば、ユーザが図示しない 操作部を操作して指示入力をおこなうことによりおこなう。
[0059] ステップ S401において、シャドウイングモードの開始要求を受け付けるまで待って( ステップ S401 : Noのループ)、受け付けた場合(ステップ S401 : Yes)、 CPU301は 、磁気ディスク 305または光ディスク 307に記録されている音源情報に対して開始ァ ドレスを設定する (ステップ S402)。音源情報は、たとえば、標準の発音で記録された 外国語の語彙や文章に関する情報である。また、開始アドレスは、たとえば、音源情 報の最初のトラックでもよいし、音源情報の再生中にユーザによりシャドウイングモー ドの開始要求がなされた位置でもよい。
[0060] そして、 CPU301は、ステップ S402において設定された開始アドレス力ら音源情 報を再生する(ステップ S403)。具体的には、 CPU301が音源情報を音声 I/F308 を介してスピーカ 310に出力する。
[0061] つぎに、 CPU301は、ステップ S402において再生中の音源情報が無音状態にな つたか否かを判断する(ステップ S404)。具体的には、 CPU301が各種センサ 316 により、再生中の音源情報の出力レベルが所定の閾値以下になったか否力、を判断 する。所定の閾値は、 ROM302記録されている音声認識プログラムにあらかじめ設 定されている。
[0062] ステップ S404において、音源情報が無音状態になっていない場合 (ステップ S404 : No)、 CPU301は、ステップ S403の処理の戻り、処理を繰り返す。一方、ステップ S 404において、音源情報が無音状態になった場合 (ステップ S404 : Yes)、 CPU301 は、再生中の音源情報を一時停止して終了アドレスを設定する(ステップ S405)。 [0063] また、 CPU301は、音声 I/F308によりスピーカ 310から案内情報を出力させる( ステップ S406)。案内情報は、開始アドレスから終了アドレスまで再生された音源情 報に相当する音声情報をユーザにマイク 309から音声で入力させる旨の情報である 。この際、 CPU301は、音声 I/F308によりマイク 309から入力される音声情報の記 録を開始する。マイク 309から入力される音声情報は、たとえば、 RAM303などに記 録される。
[0064] そして、 CPU301は、ユーザによりマイク 309から音声情報が入力されたか否かを 判断する(ステップ S407)。具体的には、各種センサ 316によりマイクから入力される 音声情報の入力レベルが所定の閾値以下の状態から所定の閾値以上の状態になつ たことを検知した場合、 CPU301は、マイク 309からユーザの音声情報の入力が開 始したと判断する。また、各種センサ 316によりマイクから入力される音声情報の入力 レベルが所定の閾値以上の状態から所定の閾値以下の状態になったことを検知した 場合、 CPU301は、マイク 309からユーザの音声情報の入力が終了したと判断する 。入力レベルの所定の閾値は、マイク 309から入力されるユーザの音声以外のノイズ に関する情報に基づいて CPU301により設定される。
[0065] ステップ S407において、マイク 309から音声情報が入力されていない場合(ステツ プ S407 : No)、 CPU301 iま、ステップ S406の処理に戻り、処理を繰り返す。一方、 ステップ S407において、マイク 309から音声情報が入力された場合(ステップ S407 : Yes)、 CPU301は、開始アドレスから終了アドレスまで再生された音源情報とマイ ク 309から入力された音声情報とを比較して差異量を算出する (ステップ S408)。具 体的には、 CPU301は、 ROM302に記録されている音声認識プログラムにしたがつ て音源情報と音声情報のスぺ外ル分析をおこなうことにより音源情報と音声情報と の差異量を算出する。この際、 CPU301は、音声情報のうちユーザの音声以外のノ ィズに関する情報を除いて音源情報との差異量を算出する。
[0066] そして、 CPU301は、各種センサ 316によりステップ S408において算出された差 異量が所定量以上であるか否力 ^判断する (ステップ S409)。差異量の所定量は、 たとえば、 RAM303などに記録されている。ステップ S409において、差異量が所定 量以上でない場合(ステップ S409 : No)、 CPU301は、ステップ S412の処理に移行 し、ステップ S405において設定した終了アドレスを開始アドレスに設定する(ステップ S412)。
[0067] 一方、ステップ S409において、差異量が所定量以上である場合(ステップ S409 :
Yes)、 CPU301は、マイク 309から入力された音声情報を削除する(ステップ S410 )。具体的には、 CPU301力 ユーザによりマイク 309から入力され、 RAM303に記 録された音声情報を削除する (ステップ S410)。
[0068] そして、 CPU301は、開始アドレスから終了アドレスまでの音源情報を所定回数再 生したか否かを判断する (ステップ S411)。ステップ S411において、音源情報を所 定回数再生していない場合(ステップ S411 : No)、 CPU301は、ステップ S403の処 理の戻り、処理を繰り返す。
[0069] 一方、ステップ S411におレ、て、音源情報を所定回数再生した場合 (ステップ S411 : Yes)、 CPU301は、ステップ S405において設定した終了アドレスを開始アドレス に設定する(ステップ S412)。そして、 CPU301は、シャドウイングモードの終了要求 を受け付けたか否かを判断する(ステップ S413)。シャドウイングモードの終了要求 は、たとえば、ユーザが図示しない操作部を操作して指示入力をおこなうことによりお こなわれる。
[0070] ステップ S413において、シャドウイングモードの終了要求を受け付けていない場合
(ステップ S413 : No)、 CPU301 iま、ステップ S403の処理に戻り、処理を繰り返す。 一方、ステップ S413において、シャドウイングモードの終了要求を受け付けた場合( ステップ S413 :Yes)、 CPU301は、一連の処理を終了する。
[0071] 以上説明したように、本実施例によれば、ナビゲーシヨン装置は、シャドウイングモ ードにおいて、再生した音源情報とユーザにマイクから入力させた音声情報との差異 量を算出する。そして、算出した差異量が所定量以上である場合は、ナビゲーシヨン 装置は、再度同じ音源情報を再生する。また、算出した差異量が所定量以下である 場合、および同じ音源情報を所定回数再生した場合、ナビゲーシヨン装置は、別の 音源情報を再生する。したがって、外国語の語彙や文章が正確に発音できないユー ザでも、シャドウイングモードを利用して外国語のあらゆる語彙や文章を発音すること により外国語の聞き取り能力の向上を図ることができる。 なお、本実施の形態で説明した情報処理方法は、あらかじめ用意されたプログラム をパーソナル.コンピュータやワークステーションなどのコンピュータで実行することに より実現することができる。このプログラムは、ハードディスク、フレキシブルディスク、 CD-ROM, M〇、 DVDなどのコンピュータで読み取り可能な記録媒体に記録され 、コンピュータによって記録媒体から読み出されることによって実行される。またこの プログラムは、インターネットなどのネットワークを介して配布することが可能な伝送媒 体であってもよい。

Claims

請求の範囲
[1] 移動体に搭載され、集音手段から入力される音声情報を処理する情報処理装置で あって、
記録媒体に記録されている音源情報を再生する再生手段と、
前記再生手段により再生されている前記音源情報を出力する出力手段と、 前記出力手段により前記音源情報が出力された後に、前記音源情報に対応する 前記音声情報を利用者に前記集音手段から入力させる入力手段と、
前記入力手段により入力された前記音声情報と前記音源情報との差異量を算出す る算出手段と、
前記算出手段により算出された前記差異量が所定量以上であるか否力を評価する 評価手段と、
前記評価手段による評価の結果、前記差異量が前記所定量以上である場合、前 記再生手段によって再生された前記音源情報を再度再生させる制御手段と、 を備えることを特徴とする情報処理装置。
[2] 前記制御手段は、前記差異量が前記所定量以上でない場合、前記再生手段によ り前記記録媒体に記録されている別の音源情報を再生させることを特徴とする請求 項 1に記載の情報処理装置。
[3] 前記制御手段は、同一の前記音源情報が所定回数再生された場合、前記再生手 段により前記記録媒体に記録されている別の音源情報を再生させることを特徴とする 請求項 1に記載の情報処理装置。
[4] 前記再生手段により再生される前記音源情報は、外国語の語彙または文章に関す る情報であり、
前記入力手段は、前記出力手段により出力された前記音源情報に相当する前記 外国語の語彙または文章を前記利用者に前記集音手段から音声で入力させることを 特徴とする請求項 1に記載の情報処理装置。
[5] 前記算出手段は、前記入力手段により入力された前記音声情報のうち前記利用者 が発した前記音声情報以外の前記音声情報を除いて前記音源情報との差異量を算 出することを特徴とする請求項:!〜 3のいずれか一つに記載の情報処理装置。
[6] 移動体に搭載され、集音手段から入力される音声情報を処理する情報処理方法で あってヽ
記録媒体に記録されている音源情報を再生する再生工程と、
前記再生工程により再生されている前記音源情報を出力する出力工程と、 前記出力工程により前記音源情報が出力された後に、前記音源情報に対応する 前記音声情報を利用者に前記集音手段から入力させる入力工程と、
前記入力工程により入力された前記音声情報と前記音源情報との差異量を算出す る算出工程と、
前記算出工程により算出された前記差異量が所定量以上であるか否力 ^評価する 評価工程と、
前記評価工程による評価の結果、前記差異量が前記所定量以上である場合、前 記再生工程によって再生された前記音源情報を再度再生させる制御工程と、 を含むことを特徴とする情報処理方法。
[7] 請求項 6に記載の情報処理方法をコンピュータに実行させることを特徴とする情報 処理プログラム。
[8] 請求項 7に記載の情報処理プログラムを記録したことを特徴とするコンピュータに読 み取り可能な記録媒体。
PCT/JP2007/054881 2006-03-20 2007-03-13 情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体 Ceased WO2007108355A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2006076878 2006-03-20
JP2006-076878 2006-03-20

Publications (1)

Publication Number Publication Date
WO2007108355A1 true WO2007108355A1 (ja) 2007-09-27

Family

ID=38522387

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2007/054881 Ceased WO2007108355A1 (ja) 2006-03-20 2007-03-13 情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体

Country Status (1)

Country Link
WO (1) WO2007108355A1 (ja)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06239186A (ja) * 1993-02-17 1994-08-30 Fujitsu Ten Ltd 車載用電子装置
JP2002175006A (ja) * 2000-12-08 2002-06-21 Ideia Corporation:Kk 語学学習装置の動作方法、語学学習装置および語学学習システム
JP2003228279A (ja) * 2002-01-31 2003-08-15 Heigen In 音声認識を用いた語学学習装置、語学学習方法及びその格納媒体
JP2004029150A (ja) * 2002-03-12 2004-01-29 Eigyotatsu Kofun Yugenkoshi コンピュータ援用の言語リスニングおよびスピーキング教示システムおよびその方法
JP2004037527A (ja) * 2002-06-28 2004-02-05 Canon Inc 情報処理装置およびその方法
JP2006201491A (ja) * 2005-01-20 2006-08-03 Advanced Telecommunication Research Institute International 発音評定装置、およびプログラム

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06239186A (ja) * 1993-02-17 1994-08-30 Fujitsu Ten Ltd 車載用電子装置
JP2002175006A (ja) * 2000-12-08 2002-06-21 Ideia Corporation:Kk 語学学習装置の動作方法、語学学習装置および語学学習システム
JP2003228279A (ja) * 2002-01-31 2003-08-15 Heigen In 音声認識を用いた語学学習装置、語学学習方法及びその格納媒体
JP2004029150A (ja) * 2002-03-12 2004-01-29 Eigyotatsu Kofun Yugenkoshi コンピュータ援用の言語リスニングおよびスピーキング教示システムおよびその方法
JP2004037527A (ja) * 2002-06-28 2004-02-05 Canon Inc 情報処理装置およびその方法
JP2006201491A (ja) * 2005-01-20 2006-08-03 Advanced Telecommunication Research Institute International 発音評定装置、およびプログラム

Similar Documents

Publication Publication Date Title
JP2644376B2 (ja) 車両用音声ナビゲーション方法
JP4804052B2 (ja) 音声認識装置、音声認識装置を備えたナビゲーション装置及び音声認識装置の音声認識方法
JP3573907B2 (ja) 音声合成装置
JP2000046577A (ja) 車両ナビゲ―ション・システムで音声案内を行うための方法および装置
CN101458093B (zh) 导航设备
JPH10149423A (ja) 地図情報提供装置
JP2002233001A (ja) 擬似エンジン音制御装置
US20090234565A1 (en) Navigation Device and Method for Receiving and Playing Sound Samples
JP2011179917A (ja) 情報記録装置、情報記録方法、情報記録プログラムおよび記録媒体
JP4030064B2 (ja) 音響経路情報を有するナビゲーションシステム
US9803991B2 (en) Route guide device and route guide method
JP6499438B2 (ja) ナビゲーション装置、ナビゲーション方法、およびプログラム
JP2023045814A (ja) ナビゲーション装置、ナビゲーション方法及びプログラム
JP2007241122A (ja) 音声認識装置、音声認識方法、音声認識プログラム、および記録媒体
CN113971892B (zh) 一种站点的播报方法、装置、多媒体设备和存储介质
WO2007108355A1 (ja) 情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体
JP2004348367A (ja) 車載情報提供装置
JP4358878B2 (ja) 音響経路情報を有するナビゲーションシステム
JP2009157065A (ja) 音声出力装置、音声出力方法、音声出力プログラムおよび記録媒体
WO2007116712A1 (ja) 音声認識装置、音声認識方法、音声認識プログラム、および記録媒体
JP2007278721A (ja) 情報処理装置、情報処理方法、情報処理プログラムおよびコンピュータに読み取り可能な記録媒体
JP2016121885A (ja) ナビゲーション装置、ナビゲーション方法、およびプログラム
JP2008157885A (ja) 情報案内装置、ナビゲーション装置、情報案内方法、ナビゲーション方法、情報案内プログラム、ナビゲーションプログラム、および記録媒体
JP4778831B2 (ja) 走行支援装置、走行支援方法、走行支援プログラムおよびコンピュータに読み取り可能な記録媒体
WO2007074739A1 (ja) データ処理装置およびデータ更新方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 07738353

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 07738353

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: JP