WO2004104986A1 - 音声出力装置及び音声出力方法 - Google Patents

音声出力装置及び音声出力方法 Download PDF

Info

Publication number
WO2004104986A1
WO2004104986A1 PCT/JP2004/006065 JP2004006065W WO2004104986A1 WO 2004104986 A1 WO2004104986 A1 WO 2004104986A1 JP 2004006065 W JP2004006065 W JP 2004006065W WO 2004104986 A1 WO2004104986 A1 WO 2004104986A1
Authority
WO
WIPO (PCT)
Prior art keywords
character
user
time
output device
delay
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2004/006065
Other languages
English (en)
French (fr)
Inventor
Makoto Nishizaki
Tomohiro Konuma
Mitsuru Endo
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Panasonic Holdings Corp
Original Assignee
Matsushita Electric Industrial Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Matsushita Electric Industrial Co Ltd filed Critical Matsushita Electric Industrial Co Ltd
Priority to US10/542,947 priority Critical patent/US7809573B2/en
Priority to JP2005504492A priority patent/JP3712207B2/ja
Publication of WO2004104986A1 publication Critical patent/WO2004104986A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/16Sound input; Sound output

Definitions

  • the present invention relates to an audio output device for conveying information by voice to a user, and more particularly to an audio output device for outputting voice and displaying characters indicating the same content as the voice.
  • an audio output device for conveying information by voice to the user
  • this audio output device is applied to terminals of a navigation system, an interface of a television, a personal computer, etc.
  • some of such voice output devices not only output voice but also display the information in characters in order to reliably convey information to the user (Japanese Patent Application Laid-Open No. Hei 1-145).
  • No. 9 5 JP-A 1 1 3 3 0 5 8 5, JP-A 2 0 0 1-1 2 4 2 4 8 4, and JP-A 5 2 1 6 6 1 8 ).
  • Even if the user misses the voice the user can grasp the information transmitted from the voice output device by reading the displayed characters without operating the voice output device.
  • FIG. 1 is a block diagram showing the configuration of a conventional voice output device that conveys information by voice and text.
  • This voice output device 900 is for interactively acquiring necessary information from the user and selling railway tickets requested by the user, and the microphone 90 1 and the voice processing unit 90 0 And 2, a transmission information generation unit 9 0 3, an audio output unit 9 0 4, and a display unit 9 0 5.
  • the microphone 901 obtains the voice from the user.
  • the voice processing unit 902 specifies, from the voice acquired by the microphone 901, user transmission information that the user intends to transmit to the audio output device 900, and transmits the user transmission information to the transmission information generation unit. Output to 9 0 3 For example, when the user issues a command to the microphone 901, the voice processing unit 902 identifies the station name "Osaka Station J" as the user transmission information.
  • the transfer information generation unit 903 generates device transfer information to be transferred to the user based on the user transfer information specified by the voice processing unit 902, and the device transfer information is generated by the voice output unit 90. 4 and output to the display unit 9 0 5. For example, when the user transmission information indicates "Osaka Station" of the departure station, the transmission information generation unit 9 0 3 generates device transmission information of the content asking for the arrival station, and transmits the device transmission information. Output.
  • the voice output unit 904 Upon acquiring the device transfer information from the transfer information generation unit 901, the voice output unit 904 outputs the contents of the device transfer information as voice. For example, when the device output information of the content asking for the arrival station is obtained, the voice output unit 940 outputs a voice “Where are you up?”.
  • the display unit 950 Upon acquiring the device transfer information from the transfer information generation unit 901, the display unit 950 displays the content of the device transfer information in characters. For example, when the display unit 9 05 acquires the device transmission information of the content asking for the arrival station, it displays the text “Where are you?”.
  • FIG. 2 is a screen display diagram showing an example of a screen displayed by the display unit 950 of the audio output device 900. -One
  • the display unit 9 05 displays the condition field 9 0 5 a, the designation field 9 0 5 b and the question field 9 0 5 c.
  • the condition column 9 05 a displays the contents to be inquired to the user such as departure station and arrival station
  • the designation column 9 0 5 b displays the station name etc. transmitted from the user
  • the question column Device transmission mentioned above to 9 0 5 c The content of the information is displayed in characters.
  • the user purchases a desired ticket by interactively operating such an audio output device 900.
  • the conventional voice output device 900 simultaneously performs voice output and character display (refer to Japanese Patent Application Laid-Open No. Hei 5-26618). For example, at the same time as the voice output unit 900 outputs the voice “Where are you?”, The display unit 900 displays the letters “Where are you?”.
  • the present invention has been made in view of the above problems, and it is an audio output device that reliably transmits information by text and voice to a user to improve the robustness of the interface with the user: Intended to provide. Disclosure of the invention
  • an audio output apparatus of the present invention there is provided a character display means for displaying transmission information to be transmitted to a user in characters; And a voice output means for outputting the transmission information as a voice when a delay time required for an operation for the user to visually recognize the character has passed after the character is displayed on the means. .
  • the voice indicating the transmission information is output. Therefore, the user moves the eyeball and focuses on the displayed character.
  • the recognition of the character and the speech recognition can be started at the same time, paying attention to both the voice and the character. As a result, text and voice information can be reliably transmitted to the user to improve the robustness of the interface with the user.
  • the voice output device further includes a delay estimation means for estimating the delay time according to the display mode of the character displayed on the character display means, and the voice output means includes the character display means for the character display means.
  • the communication information may be output as a voice when a delay time estimated by the delay estimation means has elapsed after being displayed.
  • the delay estimation unit estimates the delay time so that the delay time includes a start time until the viewpoint of the user starts moving to the character after the character is displayed by the character display unit.
  • the delay estimating means further estimates the delay time so as to include a moving time from when the user's viewpoint starts moving to reaching the character.
  • the delay estimation unit further estimates the delay time such that a focusing time from when the user's viewpoint reaches the character to when the character is focused is included.
  • the delay time is estimated according to the display mode of the character, so the user starts recognition of the character and speech recognition simultaneously. Can pay attention to both speech and letters.
  • the voice output device further includes personal information indicating a feature of the user.
  • the information processing apparatus may further include personal information acquisition means for acquiring, and the delay estimation means may estimate the delay time according to the user based on the personal information acquired by the personal information acquisition means.
  • the personal information acquisition unit acquires the age of the user as the personal information
  • the delay estimation unit determines the delay time according to the user based on the age acquired by the personal information acquisition unit.
  • the delay time is estimated based on the age that indicates the characteristics of the user, so for any individual user, delaying the voice output from the character display for the delay time according to the age of the user. Can communicate information in text and voice more reliably to the user.
  • the voice output device further includes an operation means for causing the voice output means to output a voice while displaying a character on the character display means in response to an operation by the user, and an operation means for the user to operate the operation means.
  • the delay estimation means estimates a delay time according to the user's habituation based on the degree of habituation specified by the habituation identification means. It is good. For example, the habituation specifying means specifies the number of times the user has operated the operation means as the familiarity degree.
  • the delay time is estimated based on the degree of familiarity of the user. Therefore, even if the degree of familiarity of the Rusea changes by operating the operation means, the output of voice is performed for the delay time according to the degree of familiarity It can be delayed from the text display, and text and voice information can be reliably transmitted to the user.
  • the delay estimation unit is characterized in that the focusing time is specified based on a character display distance from the fixation point of the voice output device for attracting the user's attention to the character displayed by the character display unit. As Also good.
  • the focusing time is short, and if the text display distance is long, the focusing time is long.
  • the focusing time is long.
  • the delay estimating means may estimate the delay time by using a sigmoid function.
  • the sigmoid function can represent a model of an ecosystem, by estimating the delay time using the sigmoid function in this way, it is possible to estimate an appropriate delay time that matches the biological characteristics. .
  • the delay estimation means may specify the start time based on the size of the character displayed by the character display means. Usually, the smaller the size of the character, the longer the start time, and the larger the size of the character, the shorter the start time. By thus specifying the start time based on the size of the character, an appropriate start time is identified. can do.
  • the present invention can also be realized as an audio output method in which audio is output by the audio output device, or a program thereof.
  • FIG. 1 is a block diagram showing the configuration of a conventional voice output device that conveys information by voice and text. ----
  • FIG. 2 is a screen display diagram showing an example of a screen displayed by the display unit of the same.
  • FIG. 3 is a block diagram showing the configuration of the audio output device according to the embodiment.
  • FIG. 4 is a screen display diagram showing an example of a screen displayed by the display unit of the voice output device of the same.
  • FIG. 5 is a flow chart showing the operation of the above voice output device.
  • FIG. 6 is a diagram showing the relationship between the function f 0 (X) and the motion start time T a and the character size X in the same.
  • FIG. 7 is a diagram showing the function f 1 (X) which changes with the value of the variable S in the same.
  • FIG. 8 is a diagram showing the relationship between the function f 2 (L) and the movement time T b in the same as above, and the character display distance.
  • FIG. 9 is a diagram showing the relationship between the function f 2 (L) and the focusing time T c and the character display distance L in the same.
  • FIG. 10 is a block diagram showing the configuration of the audio output device according to the first modification of the above.
  • FIG. 1 1 is a view showing the relationship between the function f 3 (M) and the individual delay time T 1 of the first modification and the age M.
  • FIG. 12 is a block diagram showing the configuration of the audio output device according to the second modification of the above.
  • FIG. 13 is a diagram showing the relationship between the function f 4 (K) and the accustomed delay time T 2 and the number of operations K in the second modification of the same.
  • FIG. 3 is a configuration diagram showing a configuration of the audio output device according to the present embodiment.
  • the voice output device 100 outputs information to be transmitted to the user as voice and displays characters indicating the information, and the microphone 1 0 1, the voice processing unit 10 2, Transmission information generator 1 0 3 and clock It comprises a unit 104, a delay unit 105, an audio output unit 106, and a display unit 107.
  • Such a voice output device 100 delays the voice output time by the time required for an operation for a person to visually recognize the letter from the letter display time (hereinafter referred to as a delay time).
  • the feature is that it makes the user recognize speech and characters surely.
  • the microphone 101 obtains voice from the user.
  • the voice processing unit 102 identifies, from the voice acquired by the microphone 1 0 1, user transmission information that the user intends to transmit to the voice output device 100, and transmits the user transmission information to the transmission information generation unit. Output to 1 0 3 For example, when the user issues a command to the microphone 101, the voice processing unit 102 identifies the station name “Osaka Station” as the user transmission information.
  • the transfer information generation unit 103 generates device transfer information to be transferred to the user based on the user transfer information specified by the voice processing unit 102, and transmits the device transfer information to the delay unit 10 5. Output to For example, when the user transfer information indicates the departure station, Osaka Station j, the transfer information generation unit 103 generates device transfer information whose contents ask for the arrival station, and outputs the device transfer information.
  • the timer unit 104 measures time in accordance with an instruction from the delay unit 105, and outputs the measurement result to the delay unit 105.
  • the delay unit 105 Upon acquiring the device transfer information from the transfer information generation unit 103, the delay unit 105 outputs the device transfer information to the display unit 100, and measures time in the clock unit 104. Start it. Then, the delay unit 105 estimates the above-mentioned delay time according to the display mode of the character displayed on the display unit 107, and the measurement time measured by the clock unit 104 reaches the delay time. At the same time, device transmission information is output to the audio output unit 106.
  • the display unit 1 0 7 obtains the device transfer information from the delay unit 1 0 5
  • the voice output unit 106 Upon acquiring the device transfer information from the delay unit 105, the voice output unit 106 outputs the content of the device transfer information as voice. For example, the voice output unit 106, upon acquiring the device transmission information of the content asking for the arrival station, outputs the voice “Where are you?”.
  • FIG. 4 is a screen display diagram showing an example of a screen displayed by the display unit 1 0 7 of the audio output device 1 00.
  • the display unit 1 0 7 has a condition field 1 0 7 a, a designation field 1 0 7 b, a question field 1 0 7 c, an agent 1 0 7 d, a start button 1 0 7 e, and a confirmation button 1 0 7 f is displayed.
  • the condition column 1 0 7 a displays the contents to be inquired of for the user, such as the departure station and the arrival station, and the designation column 1 0 7 b displays the station name etc. transmitted from the user.
  • the contents of the above-mentioned device transmission information are displayed in text on the question column 1 0 7 c.
  • the letters in question column 1 0 7 c are displayed as if agent 1 0 7 d is speaking.
  • the start point 1 0 7 e is selected by the user to start the interactive ticket sales operation of the voice output device 1 00.
  • the confirmation button 107f is selected by the user to start ticketing according to the information such as the departure station and the arrival station acquired from the user.
  • FIG. 5 is a diagram showing an operation of the audio output device 100.
  • the voice output device 100 obtains voice from the user (step S 1 0 0), and identifies user transmission information from the obtained voice (step S 1 0 2).
  • the voice output device 100 A device communication information corresponding to the information is generated (step S104), the device communication information is displayed in characters (step S106), and measurement of time is started (step S108).
  • the voice output device 100 estimates the delay time T in consideration of the character display mode, and determines whether the measurement time is equal to or longer than the delay time T. Do it (step S 1 1 0).
  • the operation from step S I 0 8 is repeatedly executed. That is, it continues to measure time.
  • the voice output device 100 determines that it is the delay time T or more (Ye s in step S 1 10), it outputs the device transmission information as voice (step S 1 1 2).
  • the delay unit 105 takes into consideration the exercise start time T a, the movement time T b, and the focusing time T c in accordance with the character display mode displayed on the display unit 1 07. Estimate the delay time T of
  • the exercise start time T a is the time required for the user's viewpoint to move toward the character after the character is displayed. For example, when the user is looking at the agent 1 0 7 d on the display unit 1 0 7 d, when the character “Where is it” is displayed in the question column 1 0 7 c, the exercise start time T a is , It is the time required for the user to remove the viewpoint from the agent 1 0 7 d which is the gaze point after the character is displayed.
  • the movement time T b is the time required for the user's viewpoint to move toward the character and to reach that character. For example, if the distance from the agent 1 0 7 d at which the user is watching to the letter in question column 1 0 7 c is long, the distance to move the viewpoint naturally is long, as a result, the moving time T b is also become longer. In such a case, the delay time T needs to be determined in consideration of the travel time Tb.
  • the focusing time T c is the time required for the user to focus on the character after the user's viewpoint has reached the character. In general, when the viewpoint is moved to see another from what a person is gazing at, the longer the movement distance, the more the focus shift occurs. Therefore, such focusing time T c is specified according to the movement distance of the observation point.
  • the exercise start time T a varies with the size of the displayed character. As the size of the character increases, the user's attention is strongly attracted to the character, and the exercise start time T a is shortened. On the other hand, as the size of the character decreases, the power to draw the user's attention to the character is weak, and the exercise start time T a becomes longer. For example, assuming that the standard size of the character is 10 points, the larger the character size than 10 points, the greater the power of attracting the user's attention to the character, and the shorter the exercise start time T a .
  • the delay unit 105 derives the exercise start time T a from the following (Expression 1).
  • T a t 0- ⁇ 0 ⁇ --(Expression "!
  • t 0 is a fixed time which is required when the size of the character is reduced as much as possible.
  • the exercise start time T a is derived by subtracting a time which varies with the size of the character from this time t 0.
  • the delay unit 105 derives a time CX O from the following (Expression 2).
  • X indicates the character size
  • t 1 is the maximum time that can be reduced by the character size X.
  • the symbol ⁇ * j indicates the product.
  • XA is a reference character size for determining exercise start time T a (eg, 10) and XC is the maximum character size (for example, 38 points) for determining the exercise start time Ta.
  • Such a function f O (X) is often used as a model of an ecosystem. That is, by using such a function f 0 (X), it is possible to derive a motion start time T a that conforms to the eye movement characteristics according to the character size X.
  • FIG. 6 is a diagram showing the relationship between the function f 0 (X) and the motion start time T a and the character size X.
  • the value indicated by the function f O (X) increases as the character size X changes from the reference size X A to the maximum point X C as shown in (a) of FIG. That is, this value gradually increases near the reference size XA (10 points) and rapidly increases around the middle size (24 points) as the character size X increases, and the maximum size XC (38 points). ) It will increase gradually again.
  • the exercise start time T a gradually decreases near the reference size XA and rapidly decreases near the middle size with the increase of the character size X, and the maximum size It gradually decreases again near XC.
  • S is a variable that determines the slope of the inflection point of the sigmoid function.
  • FIG. 7 is a diagram showing a function fl (X) which changes with the variable S.
  • the value indicated by the function f 1 (X) is a variable
  • S When S is reduced, it gradually changes around the inflection point (intermediate size), but when the variable S is increased, it changes rapidly around the inflection point.
  • this variable S By setting this variable S to an appropriate value, it is possible to derive a more accurate exercise start time Ta.
  • the movement time T b is determined by the distance from the attention point agent 1 0 7 d to the character of the question column 1 0 7 c (hereinafter referred to as character display distance).
  • the delay unit 105 derives the moving time T b from the following (Expression 5).
  • the movement time T b is derived by adding the time a 1 which changes according to the character display distance to the time t 0.
  • the delay unit 105 derives time 1 from the following (Expression 6).
  • L is the letter display distance
  • t 2 is the maximum time that can be extended by the letter display distance L.
  • L A indicates a reference distance
  • L C indicates a maximum distance.
  • the reference distance is 0 cm and the maximum distance is 1 0 cm.
  • Such a function f 2 (L) is a sigmoid function often used as a model of ecosystem. That is, by using such a function f 2 (L), it is possible to derive a moving time T b that matches the eye movement characteristics according to the character display distance. Also, the character display distance L is indicated by the following (Equation 8).
  • FIG. 8 is a diagram showing the relationship between the function f 2 (L) and the movement time T b and the character display distance L.
  • the value indicated by the function f 2 (L) increases as the character display distance changes from the reference distance L A to the maximum distance L C as shown in FIG. 8 (a). That is, this value gradually increases near the reference distance LA (O cm) with the increase of the character display distance L and rapidly increases near the intermediate distance (5 cm), and the maximum distance LC (1 O cm) It will increase gradually again.
  • the moving time T b increases gradually near the reference distance LA and rapidly increases near the intermediate distance with the increase of the distance for displaying characters. It increases moderately again near LC.
  • the character display distance L is the distance from the position of agent 1 0 7 d to the character of question column 1 0 7 c, but if agent 1 0 7 d is not displayed, the screen is displayed.
  • the distance from the center to the character may be the center of attention as the gaze point.
  • This focusing time T c is determined by the character display distance, as with the moving distance T b.
  • the delay unit 105 derives a focusing time T c from the following (Expression 9).
  • T c t 0 + ⁇ 2 ⁇ --(Expression 9)
  • t 0 is a constant time required when the character display distance L is 0.
  • the focusing time T c is derived by adding the time 2 which changes according to the character display distance to the time t 0.
  • the delay unit 1 0 5 derives the time 0 ⁇ ⁇ 2 from the following (Expression 1 0).
  • t 3 is the maximum time that can be extended by the text display distance.
  • the function f 2 (L) is given by (Eq. 7) above.
  • FIG. 9 is a diagram showing the relationship between the function f 2 (L) and the focusing time T c and the character display distance L.
  • the value indicated by the function f 2 (L) increases as the character display distance L changes from A to the maximum distance L C as shown in FIG. 9 (a).
  • the focusing time T c gradually increases near the reference distance LA with the increase of the character display distance L and rapidly increases near the intermediate distance. , It increases gradually again near the maximum distance LC.
  • the delay unit 105 derives the delay time T from the following (Expression 1 1) in consideration of the motion start time T a as described above, the movement time T b, and the focusing time T c. .
  • the delay time T is accurately determined according to the movement of the human eye by deriving the delay time T in consideration of the movement start time T a, the movement time T b, and the focusing time T c. It can be time.
  • the delay time T required for the user to visually recognize the character is estimated, and only the delay time T elapses after the character is displayed.
  • the user can also start speech recognition as well as start character recognition. As a result, it is possible to reliably convey information in the form of text and voice to the user and to improve the robustness of the interface between the user and the user.
  • the time ⁇ 1 and the time 2 are displayed regardless of the character display distance L.
  • Each may be a fixed time. That is, let time 1 and time 2 be average times that can be taken according to the change of the character display distance L.
  • the delay unit 105 derives the delay time ⁇ ⁇ from the following (Expression 1 2).
  • the averaging time By using the averaging time in this manner, the number of parameters for deriving the delay time ⁇ can be reduced, and the calculation process can be simplified. As a result, the calculation speed of the delay time ⁇ can be increased, and the configuration of the delay unit 105 can be simplified.
  • the voice output device estimates a delay time according to each user, and more specifically, estimates according to the age of each user.
  • FIG. 10 is a block diagram showing the configuration of the audio output device according to the first modification.
  • the voice output device 1 OO a includes a microphone 1 0 1, a voice processing unit 1 2 0 2, a transmission information generation unit 1 0 3, a clocking unit 1 0 4, and a delay unit 1 0 5 a And an audio output unit 106, a display unit 1 0 7, a card reader 1 0 9, and a personal information storage unit 1 0 8.
  • the card reader 100 reads out the age and date of birth, which is personal information, from the card 1 0 9a inserted into the voice output device 1 OO a by the user, and reads out the read personal information in the personal information storage unit 1 0 Store in 8
  • the delay unit 105a derives the delay time T in consideration of the motion start time T a, the movement time T b, and the focusing time T c as described above. Then, the delay unit 105 a refers to the personal information stored in the personal information storage unit 108 and derives from the delay time T an individual delay time T 1 in consideration of the personal information. Furthermore, the delay unit 105a outputs the device transfer information as voice from the voice output unit 106, and then displays the device transfer information after the elapse of the individual delay time T 1. Make it appear in characters.
  • the delay unit 105 derives the individual delay time T 1 from the following (Expression 13).
  • T 1 T + ⁇ 3 ⁇ --(Expression 1 3)
  • the individual delay time ⁇ 1 is derived by adding the time 3 that changes with age to the delay time ⁇ .
  • the delay unit 105 derives a time ⁇ 3 from the following (Expression 14).
  • t 4 is the maximum time that can be extended by the age M Ru.
  • Figure 1 1 shows the relationship between the function f 3 (M) and the individual delay time T 1 and the age M.
  • the value indicated by the function f 3 (M) increases as the age M changes from 20 years to 60 years, as shown in (a) of FIG. That is, with age, this value gradually increases around 20 years of active exercise (base age), rapidly increases around 40 years of middle age, and exercise ability declines. It gradually increases again near (maximum age).
  • the individual delay time T 1 is derived in consideration of the age of the user, and the voice is output after a lapse of the individual delay time T 1 after the character is displayed.
  • Each user can improve the robustness of the interface: c-c.
  • the age of the user is used as personal information, but the reaction speed of the user, the movement speed of the viewpoint, the focusing speed, the agility, the usage history, and the like may be used.
  • personal information such as the reaction rate as described above is pre-registered in the card 1 0 9 a, and the card reader 1 0 9 reads the personal information from the card 1 0 9 a and stores the personal information.
  • the delay unit 1 0 5 a is stored in the personal information storage unit 1 0 8
  • personal information such as reaction rate
  • individual delay time ⁇ 1 is derived from delay time T considering its reaction rate.
  • the voice output device is to estimate the delay time according to the user's familiarity, and more specifically to estimate according to the number of times of user operation.
  • the movement start time T a and the movement time T b and the focusing time T c become shorter because the user gets used to the operation.
  • the voice output device estimates the delay time in consideration of the number of user operations.
  • FIG. 12 is a block diagram showing the configuration of the audio output device according to the second modification.
  • the voice output device 1 OO b according to the second modification includes a microphone 1 0 1, a voice processing unit 1 0 2, a transmission information generation unit 1 0 3, a clocking unit 1 0 4, and a delay unit 1 0 5 b , An audio output unit 1 0 6, a display unit 1 0 7, and a counter 1 1 0.
  • the counter 1 1 0 When the counter 1 1 0 obtains the user transmitted information output from the voice processing unit 1 0 2, it counts the number of times of obtaining the number and the number of operations on the voice output device 1 0 O b of the false alarm, and delays the number of operations. Notify part 1 0 5 b.
  • the delay unit 105 b first derives the delay time T in consideration of the motion start time T a, the movement time T b, and the focusing time T c as described above.
  • the delay unit 1 0 5 b refers to the number of operations notified from the counter 1 1 0 and derives from the delay time T a habituation delay time T 2 in consideration of the number of operations.
  • the delay unit 105b causes the device output information to be output as voice from the voice output unit 106, the device transfer information is displayed on the display unit 1 07 after elapse of the familiar delay time T2. Display in characters.
  • the delay unit 105 b derives the accustomed delay time T 2 from the following (Expression 16).
  • T 2 T-Of 4- ⁇ '(equation 1 6)
  • the accustomed delay time T2 is derived by subtracting the time 4 which changes according to the number of operations from the delay time.
  • the delay unit 105 b derives the time 0? 4 from the following (Expression 1 7).
  • is the number of operations
  • t 5 is the maximum time that can be reduced by the number of operations K.
  • K C is the maximum value of the number of operations such that the accustomed delay time T 2 is the shortest.
  • FIG. 13 is a diagram showing the relationship between the function f 4 (K) and the familiar delay time T 2 and the number of operations K.
  • the value indicated by the function f 4 (K) increases as the number of operations K changes from 0 (reference number) to KC (maximum number) as shown in (a) of Fig. 13. That is, this value gradually increases near 0 times of the completely unfamiliar state with the increase of the operation number K, and KC 2 It rapidly increases around the middle (intermediate number) and gradually increases again near the fully used KC (maximum number).
  • the accustomed delay time T 2 decreases gradually near the reference number with the increase of the operation number K and decreases rapidly near the middle number, It decreases gradually again around the maximum number of times.
  • the familiar delay time T 2 is derived in consideration of the user's familiarity, and after displaying the characters, only the familiar delay time T 2 elapses, and then the voice is output.
  • the robustness of the interface can be maintained.
  • the audio output device estimates the delay time according to the user's familiarity, as in the second modification, and specifically estimates according to the user's operation time.
  • the voice output device estimates the delay time in consideration of the user's operation time.
  • the audio output device has the same configuration as the audio output device 1 O O b of the second modification shown in FIG. 12, but the operations of the delay unit 1 0 5 b and the counter 1 1 0 are different.
  • the counter 1 1 0 has a function as a time counter, and after the dialog between the voice output device 1 0 0 b and the user is started, the first user from the voice processing unit 1 0 2 When transmission information is acquired, the elapsed time from the acquisition, that is, the operation time, is measured. And counter 1 1 0 notifies the delay time to the delay unit 1 0 5 b.
  • the delay unit 105 b first derives the delay time T in consideration of the motion start time T a, the movement time T b, and the focusing time T c as described above. Then, the delay unit 1 0 5 b refers to the operation time notified from the counter 1 1 0, and derives from the delay time T a habituation delay time T 3 in consideration of the operation time. Furthermore, the delay unit 105 b causes the device output information to be output as voice from the voice output unit 106, and the device transfer information is displayed on the display unit 1 07 after elapse of the accustomed delay time T 3. Display in characters.
  • the delay unit 105 b derives the accustomed delay time T 3 from the following (Expression 19).
  • T 3 T- ⁇ 5 ⁇ --(Expression 1 9)
  • the accustomed delay time T3 is derived by subtracting the time 5 which changes according to the operation time from the delay time.
  • the delay unit 105 b derives a time ⁇ 5 from the following (Eq. 20).
  • t 6 is the maximum time that can be reduced by the operation time P.
  • P C is the maximum value of the operation time P such that the accustomed delay time T 3 is the shortest.
  • the familiar delay time T3 is derived in consideration of the user's familiarity, and after displaying the characters, the voice is output after the familiar delay time T3 has elapsed. Because of this, it is You can maintain the robustness of the face.
  • the measurement of the operation time P is started at the timing when the voice processing unit 102 outputs the user transmission information, that is, the timing when the user utters the voice.
  • the measurement of the operation time P may be started at the timing when the is input or when the start button 1 0 7 f is selected.
  • the exercise start time T a changes with not only the size of the character but also with the display position of the character. That is, as the position of the displayed character is closer to the user's gaze point, the user notices the character earlier, so the exercise start time T a becomes shorter.
  • the delay unit 105 derives the motion start time T a from the following (Expression 2 2) based on the character display distance.
  • the exercise start time T a is derived by adding a time 6 which changes according to the character display distance to the time t 0.
  • the delay unit 105 derives 6 from the following (Expression 2 3).
  • t 7 is the maximum time that can be extended by the character display distance. Also, the function f 2 (L) is shown by (Eq. 7).
  • a delay time T based on the exercise start time T a in consideration of the character display distance L is derived, and voice is output after the delay time T has elapsed after displaying the character.
  • An interface suitable for each character display distance L The robustness of the face can be maintained.
  • the delay unit 105 derives the exercise start time T a from the following (Equation 24) based on the fixation point and the contrast of the characters.
  • the motion start time T a is derived by subtracting the time ⁇ 7, which varies depending on the contrast, from this time t 0.
  • the delay unit 105 derives 7 from the following (Expression 25).
  • t 8 is the maximum time that can be shortened by the contrast Q.
  • QA is the reference contrast for determining the exercise start time Ta
  • QC is the maximum contrast for determining the exercise start time Ta.
  • the delay time T is derived based on the movement start time T a in consideration of the contrast, and the voice is output after the delay time T has elapsed since the character is displayed. It is possible to maintain the robustness of the interface suitable for each contrast. (Modification 6)
  • the delay unit 105 derives the exercise start time T a from the following (Expression 27) based on the degree of emphasis of the character display mode.
  • t 0 is a fixed time which is required when the degree of emphasis of the display mode is reduced without limit. That is, the exercise start time T a is derived by subtracting the time 8 which changes according to the degree of emphasis from this time t O.
  • the delay unit 105 derives QT 8 from the following (Expression 2 8).
  • R indicates the degree of emphasis
  • t 9 is the maximum time that can be shortened by the degree of emphasis R.
  • R A is the emphasis degree of the criteria for determining the exercise start time T a
  • R C is the maximum emphasis degree for determining the exercise start time T a.
  • the delay time T based on the exercise start time T a in consideration of the degree of emphasis of the character is derived, and after displaying the character, the voice is output after the delay time T has elapsed.
  • the robustness of the interface can be maintained for each degree of emphasis on
  • the motion start time T a, the movement time T b, and the focusing time T c are all considered in order to derive the delay time T, but at least one of them is taken into consideration.
  • the delay time may be derived in consideration of one.
  • the delay time is derived based on the personal information of the user
  • the delay time is derived based on the familiarity of the user, but the delay is calculated based on both the personal information and familiarity of the user
  • the time may be derived.
  • the voice output device is described as a device for selling tickets by voice output and character display. However, the voice output device and character display can be used. If there is, it may be a device that performs other operations. For example, a TV, a terminal of a car navigation system, a mobile phone, a mobile phone, a personal computer, a telephone, a fax machine, a microwave, a refrigerator, a vacuum cleaner, an electronic dictionary, an electronic translator, etc. May be configured.
  • the voice output device can reliably convey information in the form of text and voice to the user to improve the robustness of the interface with the user. It is useful for voice response devices that sell tickets etc. by responding with.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • User Interface Of Digital Computer (AREA)
  • Electrically Operated Instructional Devices (AREA)
  • Ticket-Dispensing Machines (AREA)

Abstract

文字と音声による情報をユーザに対して確実に伝えてユーザとの間のインターフェースの頑健性を向上する音声出力装置は、ユーザに対して伝達すべき装置伝達情報を文字で表示する表示部(107)と、ユーザがその表示部(107)で表示される文字を視認するための動作に要する遅延時間(T)を推定し、その文字が表示されてから遅延時間(T)が経過したときに、その装置伝達情報を音声で出力する遅延部(105)及び音声出力部(106)とを備える。

Description

明 細 書 音声出力装置及び音声出力方法 技術分野
本発明は、 ュ一ザに対して音声で情報を伝える音声出力装置に関し、 特に、 音声を出力するとともにその音声と同一の内容を示す文字を表示 する音声出力装置に関する。 背景技術
従来より、 ユーザに対して音声で情報を伝える音声出力装置が提供さ れ、 この音声出力装置は、 力一ナビゲ一シヨンシステムの端末や、 テレ ビ、パーソナルコンピュータなどのインタ一フェースに適用されている。 また、 このような音声出力装置には、 ユーザに対して情報を確実に伝 えるため、 音声を出力するのみならず、 その情報を文字で表示するもの もある (特開平 1 1一 1 4 5 9 5 5号公報、 特開平 1 1 — 3 3 9 0 5 8 号公報、 特開 2 0 0 1 — 1 4 2 4 8 4号公報、 及び特開平 5— 2 1 6 6 1 8号公報参照) 。 仮にユーザが音声を聞き逃しても、 ユーザは音声出 力装置をわざわざ操作することなく、 表示された文字を読むことで音声 出力装置から伝えられる情報を把握することができる。
図 1 は、 音声及び文字で情報を伝える従来の音声出力装置の構成を示 す構成図である。
この音声出力装置 9 0 0は、 対話形式でユーザから必要な情報を取得 して、 そのユーザの要望する鉄道の切符を販売するものであって、 マイ ク 9 0 1 と、 音声処理部 9 0 2と、 伝達情報生成部 9 0 3 と、 音声出力 部 9 0 4と、 表示部 9 0 5とを備える。 マイク 9 0 1 は、 ユーザからの音声を取得する。
音声処理部 9 0 2は、 マイク 9 0 1 で取得された音声から、 ユーザが 音声出力装置 9 0 0に対して伝えようとするユーザ伝達情報を特定し、 そのユーザ伝達情報を伝達情報生成部 9 0 3に出力する。 例えば、 ュ一 ザがマイク 9 0 1 に対して Γォォサ力」 と発すると、 音声処理部 9 0 2 は、 駅名の 「大阪駅 J をユーザ伝達情報として特定する。
伝達情報生成部 9 0 3は、 音声処理部 9 0 2で特定されたユーザ伝達 情報に基づいて、 ユーザに対して伝えるべき装置伝達情報を生成し、 そ の装置伝達情報を音声出力部 9 0 4及び表示部 9 0 5に出力する。 例え ぱ、 ユーザ伝達情報が出発駅の 「大阪駅」 を示す場合には、 伝達情報生 成部 9 0 3は、 到着駅を尋ねる内容の装置伝達情報を生成して、 その装 置伝達情報を出力する。
音声出力部 9 0 4は、 伝達情報生成部 9 0 3から装置伝達情報を取得 すると、 その装置伝達情報の内容を音声で出力する。 例えば、 音声出力 部 9 0 4は、 到着駅を尋ねる内容の装置伝達情報を取得すると、 「どこ までですか」 という音声を出力する。
表示部 9 0 5は、 伝達情報生成部 9 0 3から装置伝達情報を取得する と、 その装置伝達情報の内容を文字で表示する。 例えば、 表示部 9 0 5 は、 到着駅を尋ねる内容の装置伝達情報を取得すると、 「どこまでです か」 という文字を表示する。
図 2は、 音声出力装置 9 0 0の表示部 9 0 5が表示する画面の一例を 示す画面表示図である。 - 一
表示部 9 0 5は、 条件欄 9 0 5 a と指定欄 9 0 5 bと質問欄 9 0 5 c とを表示する。 条件欄 9 0 5 aには、 ユーザに対して出発駅や到着駅な どの問い合わせるべき内容が表示され、 指定欄 9 0 5 bには、 ユーザか ら伝えられた駅名などが表示され、 質問欄 9 0 5 cには上述の装置伝達 情報の内容が文字で表示される。
ユーザは、 このような音声出力装置 9 0 0を対話形式で操作すること によリ所望の切符を購入する。
ここで、 従来の音声出力装置 9 0 0は、 音声の出力と文字の表示とを 同時に行う (特開平 5— 2 1 6 6 1 8号公報参照) 。 例えば、 音声出力 部 9 0 4が 「どこまでですか」 という音声を出力すると同時に、 表示部 9 0 5が 「どこまでですか」 という文字を表示する。
しかしながら、 上記従来の音声出力装置 9 0 0では、 音声が文字の表 示と同時に出力されるため、 ユーザの注意力が文字よりも音声に集中し てしまって、 文字表示がユーザにとって無意味となり、 ユーザとの間で のイ ンターフェースの頑健性を向上することができないという問題があ る。
これは、 人は文字が表示されてもその文字を認識するのに時間を要し てしまうからである。 人は、 文字が表示されてから眼球運動を開始する までに、 7 0 m sから 7 0 0 m sま の時間を要することが知られてい る (田村博著 「ヒューマンインターフェース」 ( 1 9 9 8年オーム社発 行) 参照) 。 また、 その時間の平均は 2 0 0 m sである。 更に、 視点を 文字の位置まで動かし、 焦点をその文字に合わせるまでには、 それ以上 の時間が必要とされる。
本発明は、 かかる問題に鑑みてなされたものであり、 文字と音声によ る情報をユーザに対して確実に伝えてユーザとの間のインターフ: n—ス の頑健性を向上する音声出力装置を提供することを目的とする。 発明の開示
上記目的を達成するために、 本発明の音声出力装置は、 ユーザに対し て伝達すべき伝達情報を文字で表示する文字表示手段と、 前記文字表示 手段に文字が表示されてから、 ユーザが前記文字を視認するための動作 に要する遅延時間が経過したときに、 前記伝達情報を音声で出力する音 声出力手段とを備えることを特徴とする。 .
これにより、 伝達情報を示す文字が表示されてから遅延時間経過後 ίこ その伝達情報を示す音声が出力されるため、 ユーザは、 眼球を動かして 表示された文字に焦点を合わせた状態から、 その文字の認識と、 音声の 認識とを同時に開始し、 音声と文字の両方に注意を払うことができる。 その結果、 文字と音声による情報をユーザに対して確実に伝えてユーザ との間のインターフェースの頑健性を向上することができる。
また、 前記音声出力装置は、 さらに、 前記文字表示手段に表示される 文字の表示態様に応じて前記遅延時間を推定する遅延推定手段を備え、 前記音声出力手段は、 前記文字表示手段に文字が表示されてから、 前記 遅延推定手段によリ推定された遅延時間が経過したときに、 前記伝達情 報を音声で出力することを特徴としても良い。 例えば、 前記遅延推定手 段は、 前記文字表示手段により文字が表示されてから、 ユーザの視点が 前記文字に移動を開始するまでの開始時間を含むように前記遅延時間を 推定する。 また、 前記遅延推定手段は、 さらに、 ユーザの視点が移動を 開始してから前記文字に到達するまでの移動時間を含むように前記遅延 時間を推定する。 また、 前記遅延推定手段は、 さらに、 ユーザの視点が 前記文字に到達してから前記文字に焦点が合うまでの焦点合わせ時間を 含むように前記遅延時間を推定する。
これにより、文字表示手段に表示される文字の表示態様が異なっても、 その文字の表示態様に応じて遅延時間が推定されるため、 ユーザはその 文字の認識と音声の認識とを同時に開始して、 音声と文字の両方に注意 を払うことができる。
また、 前記音声出力装置は、 さらに、 ユーザの特徴を示す個人情報を 取得する個人情報取得手段を備え、 前記遅延推定手段は、 前記個人情報 取得手段により取得された個人情報に基づいて、 前記ユーザに応じた遅 延時間を推定することを特徴と しても良い。 例えば、 前記個人情報取得 手段は、 前記ユーザの年齢を前記個人情報と して取得し、 前記遅延推定 手段は、 前記個人情報取得手段により取得された年齢に基づいて、 前記 ユーザに応じた遅延時間を推定する。
これにより、 遅延時間がユーザの特徴を示す年齢に基づいて推定され るため、 個々のユーザの何れに対しても、 そのユーザの年齢に応じた遅 延時間だけ音声の出力を文字表示から遅らせることができ、 文字と音声 による情報をユーザに対してより確実に伝えることができる。
また、 前記音声出力装置は、 さらに、 ユーザによる操作に応じて、 前 記文字表示手段に文字を表示させるとともに、 前記音声出力手段に音声 を出力させる操作手段と、 ユーザの前記操作手段に対する操作の慣れ度 合いを特定する慣れ特定手段とを備え、 前記遅延推定手段は、 前記慣れ 特定手段により特定された慣れ度合いに基づいて、 前記ユーザの慣れに 応じた遅延時間を推定することを特徴と しても良い。 例えば、 前記慣れ 特定手段は、 前記操作手段に対するユーザの操作回数を前記慣れ度合い として特定する。
これにより、 遅延時間がユーザの慣れ度合いに基づいて推定されるた め、 操作手段を操作することによリューザの慣れ度合いが変化しても、 その慣れ度合いに応じた遅延時間だけ音声の出力を文字表示から遅らせ ることができ、 文字と音声による情報をユーザに対してよリ確実に伝え ることができる。
また、 前記遅延推定手段は、 ユーザの注意を引き付ける前記音声出力 装置の注視点から、 前記文字表示手段により表示された文字までの文字 表示距離に基づいて、 前記焦点合わせ時間を特定することを特徴として も良い。
通常、 文字表示距離が短ければ焦点合わせ時間も短く、 文字表示距離 が長ければ焦点合わせ時間も長いため、 このように焦点合わせ時間を文 字表示距離に基づいて特定することにより、 適切な焦点合わせ時間を特 定することができる。
また、 前記遅延推定手段は、 シグモイ ド関数を利用して前記遅延時間 を推定することを特徴と しても良い。
シグモイ ド関数は生態系のモデルを表現し得るものであるため、 この ようにシグモイ ド関数を利用して遅延時間を推定することにより、 生体 特性に合致した適切な遅延時間を推定することができる。
また、 前記遅延推定手段は、 前記文字表示手段により表示される文字 のサイズに基づいて、前記開始時間を特定することを特徴としても良い。 通常、 文字のサイズが小さければ開始時間は長く、 文字のサイズが大 きければ開始時間は短くなるため、 このように開始時間を文字のサイズ に基づいて特定することにより、 適切な開始時間を特定することができ - る。
なお、 本発明は、 上記音声出力装置によって音声が出力される音声出 力方法やそのプログラムと して実現することもできる。 図面の簡単な説明
図 1 は、 音声及び文字で情報を伝える従来の音声出力装置の構成を示 す構成図である。 - ……
図 2は、同上の表示部が表示する画面の一例を示す画面表示図である。 図 3は、実施の形態における音声出力装置の構成を示す構成図である。 図 4は、 同上の音声出力装置の表示部が表示する画面の一例を示す画 面表示図である。 図 5は、 同上の音声出力装置の動作を示すフロー図である。
図 6は、 同上の関数 f 0 ( X ) 及び運動開始時間 T a と文字サイズ X との関係を示す図である。
図 7は、 同上の変数 Sの値によって変化する関数 f 1 ( X ) を示す図 である。
図 8は、 同上の関数 f 2 ( L ) 及び移動時間 T bと文字表示距離しと の関係を示す図である。
図 9は、 同上の関数 f 2 ( L ) 及び焦点合わせ時間 T cと文字表示距 離 Lとの関係を示す図である。
図 1 0は、 同上の変形例 1 にかかる音声出力装置の構成を示す構成図 である。
図 1 1 は、 同上の変形例 1 の関数 f 3 ( M ) 及び個人別遅延時間 T 1 と年齢 Mとの関係を示す図である。
図 1 2は、 同上の変形例 2にかかる音声出力装置の構成を示す構成図 である。
図 1 3は、 同上の変形例 2の関数 f 4 ( K ) 及び慣れ遅延時間 T 2と 操作回数 Kとの関係を示す図である。 発明を実施するための最良の形態
以下、 本発明の実施の形態における音声出力装置について、 図面を参 照しながら説明する。
図 3は、 本実施の形態における音声出力装置の構成を示す構成図であ る。
本実施の形態における音声出力装置 1 0 0は、 ユーザに伝える情報を 音声で出力するとともにその情報を示す文字を表示するものであって、 マイク 1 0 1 と、 音声処理部 1 0 2と、 伝達情報生成部 1 0 3と、 計時 部 1 0 4と、 遅延部 1 0 5と、 音声出力部 1 0 6と、 表示部 1 0 7とを 備えている。
このような音声出力装置 1 0 0は、 音声の出力時刻を、 文字の表示時 刻から、 人がその文字を視認するための動作に要する時間 (以下、 遅延 時間という) だけ遅らせることにより、 その音声と文字とをユーザに対 して確実に認識させる点に特徴がある。
マイク 1 0 1 は、 ユーザからの音声を取得する。
音声処理部 1 0 2は、 マイク 1 0 1 で取得された音声から、 ユーザが 音声出力装置 1 0 0に対して伝えようとするユーザ伝達情報を特定し、 そのユーザ伝達情報を伝達情報生成部 1 0 3に出力する。 例えば、 ュ一 ザがマイク 1 0 1 に対して Γォォサ力」 と発すると、 音声処理部 1 0 2 は、 駅名の 「大阪駅」 をユーザ伝達情報と して特定する。
伝達情報生成部 1 0 3は、 音声処理部 1 0 2で特定されたユーザ伝達 情報に基づいて、 ユーザに対して伝えるべき装置伝達情報を生成し、 そ の装置伝達情報を遅延部 1 0 5に出力する。 例えば、 ユーザ伝達情報が 出発駅の Γ大阪駅 j を示す場合には、 伝達情報生成部 1 0 3は、 到着駅 を尋ねる内容の装置伝達情報を生成して、その装置伝達情報を出力する。 計時部 1 0 4は、 遅延部 1 0 5からの指示に応じて時間を計測し、 そ の計測結果を遅延部 1 0 5に出力する。
遅延部 1 0 5は、 伝達情報生成部 1 0 3から装置伝達情報を取得する と、 その装置伝達情報を表示部 1 0 7に対して出力するとともに、 計時 部 1 0 4に時間の計測を開始させる。 そして、 遅延部 1 0 5は、 表示部 1 0 7に表示される文字の表示態様に応じて上述の遅延時間を推定し、 計時部 1 0 4で計測された計測時間が遅延時間に達したときに、 装置伝 達情報を音声出力部 1 0 6に出力する。
表示部 1 0 7は、 遅延部 1 0 5から装置伝達情報を取得すると、 その 装置伝達情報の内容を文字で表示する。 例えば、 表示部 1 0 7は、 到着 駅を尋ねる内容の装置伝達情報を取得すると、 「どこまでですか J とい う文字を表示する。
音声出力部 1 0 6は、 遅延部 1 0 5から装置伝達情報を取得すると、 その装置伝達情報の内容を音声で出力する。 例えば、 音声出力部 1 0 6 は、 到着駅を尋ねる内容の装置伝達情報を取得すると、 「どこまでです か」 という音声を出力する。
図 4は、 音声出力装置 1 0 0の表示部 1 0 7が表示する画面の一例を 示す画面表示図である。
表示部 1 0 7は、 条件欄 1 0 7 a と、 指定欄 1 0 7 bと、 質問欄 1 0 7 c と、 エージェン ト 1 0 7 d と、 開始ボタン 1 0 7 e と、 確認ポタン 1 0 7 f とを表示する。
条件欄 1 0 7 aには、 ユーザに対して出発駅や到着駅などの問い合わ せるべき内容が表示され、 指定欄 1 0 7 bには、 ュ一ザから伝えられた 駅名などが表示され、 質問欄 1 0 7 cには上述の装置伝達情報の内容が 文字で表示される。 また、 質問欄 1 0 7 cの文字は、 あたかもエージェ ン ト 1 0 7 dが喋っているように表示される。
開始ポタン 1 0 7 eは、 ユーザに選択されることにより、 音声出力装 置 1 0 0の対話形式による切符販売動作を開始させる。
確認ボタン 1 0 7 f は、 ユーザに選択されることにより、 ユーザから 取得された出発駅や到着駅などの情報に応じた切符の発券を開始する。 図 5は、 音声出力装置 1 0 0の動作を示すプロ一図である。
まず、 音声出力装置 1 0 0は、 ユーザからの音声を取得し (ステップ S 1 0 0 ) 、 その取得した音声からユーザ伝達情報を特定する (ス亍ッ プ S 1 0 2 ) 。
次に、 音声出力装置 1 0 0は、 そのユーザ伝達情報に基づいて、 その 情報に対応する装置伝達情報を生成し (ステップ S 1 0 4 ) 、 その装置 伝達情報を文字で表示するとともに (ステップ S 1 0 6 ) 、 時間の計測 を開始する (ステップ S 1 0 8 ) 。
このように時間の計測を開始すると、 音声出力装置 1 0 0は、 その文 字の表示態様を考慮して遅延時間 Tを推定し、 計測時間がその遅延時間 T以上であるか否かを判別する (ステップ S 1 1 0 ) 。 ここで、 音声出 力装置 1 0 0は、 遅延時間 T未満であると判別すると (ステップ S 1 1 0の N o ) 、 ステップ S I 0 8からの動作を繰り返し実行する。 即ち、 時間計測を継続して実行する。 一方、 音声出力装置 1 0 0は、 遅延時間 T以上であると判別すると (ステップ S 1 1 0の Y e s ) 、 装置伝達情 報を音声で出力する (ステップ S 1 1 2 ) 。
ここで、 遅延部 1 0 5は、 表示部 1 0 7に表示される文字の表示態様 に応じ、 運動開始時間 T a と移動時間 T bと焦点合わせ時間 T c とを考 慮して、 上述の遅延時間 Tを推定する。
運動開始時間 T aは、 文字が表示されてから、 ユーザの視点がその文 字に向かって動き出すまでに要する時間である。 例えば、 ユーザが表示 部 1 0 7のエージェン ト 1 0 7 dを注視しているときに、 「どこまでで すか」 という文字が質問欄 1 0 7 cに表示される場合、 運動開始時間 T aは、 その文字が表示されてから、 ユーザが注視点であるエージェント 1 0 7 dから視点を外すまでに要する時間である。
移動時間 T bは、 ユーザの視点が文字に向かって動き出してから、 そ の文字に到達するまでに要する時間である。 例えば、 ユーザが注視して いるエージェン ト 1 0 7 dから、 質問欄 1 0 7 cの文字までの距離が長 いと、 当然に視点を動かす距離も長くなリ、 その結果、 移動時間 T bも 長くなる。 このような場合には、 遅延時間 Tはその移動時間 T bを考慮 して決定される必要がある。 焦点合わせ時間 T cは、 ユーザの視点が文字に到達してから、 その文 字に焦点が合うまでに要する時間である。 一般に、 人が注視しているも のから別のものを見るために視点を移動すると、 その移動距離が長いほ ど焦点のずれが生じる。 そこで、 このような焦点合わせ時間 T cは、 視 点の移動距離に応じて特定される。
ここで、 運動開始時間 T aについて詳細に説明する。 '
この運動開始時間 T aは、表示される文字のサイズによって変化する。 文字のサイズが大きくなれば、 ユーザの注意はその文字に強く引き付け られて、 運動開始時間 T aは短くなる。 一方、 文字のサイズが小さくな れぱ、 ユーザの注意を文字に引き付ける力は弱く、 運動開始時間 T aは 長くなる。 例えば、 文字の標準のサイズを 1 0ポイン トとすると、 文字 のサイズを 1 0ポイン トより大きくするほど、 ユーザの注意を文字に引 き付ける力が大きくなり、 運動開始時間 T aは短くなる。
遅延部 1 0 5は、 以下の (式 1 ) から運動開始時間 T aを導出する。 T a = t 0 - α 0 ■ - - (式 "! )
t 0は、 文字のサイズを限りなく小さく.したときに要する一定の時間 である。 運動開始時間 T aは、 この時間 t 0から、 文字のサイズによつ て変化する時間ひ 0を減算することによって導出される。
遅延部 1 0 5は、 以下の (式 2 ) から時間 CX Oを導出する。
α 0 = t 1 * f 0 ( X ) ■ ' ' (式 2 )
Xは文字サイズを示し、 t 1 は文字サイズ Xによって短縮可能な最大 の時間である。 なお、 記号 Γ * j は積を示す。
また、 関数 f O ( X ) は以下の (式 3 ) によって示される。
f 0 ( X ) = 1/(1+exp (- ((X-XA)/(XC-XA) -0.5)/0.1) ) ' ■ ■ (式 3 )
X Aは運動開始時間 T a を決めるための基準の文字サイズ (例えば、 1 0ポイント) であり、 X Cは運動開始時間 T aを決めるための最大の 文字サイズ (例えば、 3 8ポイン ト) である。 なお、 記号 「 e X p J は 自然対数の底を示し、 e x p (A ) は、 自然対数の底の A乗を示す。 このような関数 f O ( X ) は、 生態系のモデルとして良く利用される シグモイ ド関数である。 即ち、 このような関数 f 0 ( X ) を用いること で、 文字サイズ Xに応じた眼球運動特性に適合する運動開始時間 T aを 導出することができる。
図 6は、 関数 f 0 ( X ) 及び運動開始時間 T a と文字サイズ Xとの関 係を示す図である。
関数 f O ( X ) によって示される値は、 図 6の ( a ) に示すように、 文字サイズ Xが基準サイズ X Aから最大ポイン ト X Cに変化するに伴つ て増加する。 即ち、 この値は文字サイズ Xの増加とともに、 基準サイズ X A ( 1 0ポイン ト) 付近で緩やかに増加し、 中間サイズ ( 2 4ポイン ト) 付近で急激に増加し、 最大サイズ X C ( 3 8ポイント) 付近で再び 緩やかに増加する。
したがって、 運動開始時間 T aは、 図 6の ( b ) に示すように、 文字 サイズ Xの増加に伴って、 基準サイズ X A付近で緩やかに減少し、 中間 サイズ付近で急激に減少し、 最大サイズ X C付近で再び緩やかに減少す る。
ここで、 関数 f 0 ( X ) の代わりに (式 4 ) に示す関数を用いても良 い。
f 1 ( X ) = 1/(1+exp(-S*((X-XA)/(XG-XA) -0.5)/0.1)) ■ · ' (式
4 )
Sはシグモイ ド関数の変極点の傾きを決定する変数である。
図 7は、 変数 Sによって変化する関数 f l ( X ) を示す図である。 この図 7に示すように、 関数 f 1 ( X ) によって示される値は、 変数 Sを小さくすると、 変極点 (中間サイズ) 付近で緩やかに変化するが、 変数 Sを大きくすると、 変極点付近で急激に変化する。 この変数 Sを適 当な値に設定することにより、 より正確な運動開始時間 T aを導出する ことができる。
次に、 移動時間 T bについて詳細に説明する。
この移動時間 T bは、 注視点であるエージェン ト 1 0 7 dから質問欄 1 0 7 cの文字までの距離 (以下、 文字表示距離という) によって定め られる。
遅延部 1 0 5は、 以下の (式 5 ) から移動時間 T bを導出する。
T b = t 0 + α 1 - - - (式 5 )
t 0は、 文字表示距離が 0のときに要する一定の時間である。 移動時 間 T bは、 この時間 t 0に対して、 文字表示距離に応じて変化する時間 a 1 を加算することによって導出される。
遅延部 1 0 5は、 以下の (式 6 ) から時間 1 を導出する。
α 1 = t 2 * f 2 ( L ) ■ ■ ■ (式 6 )
Lは文字表示距離であり、 t 2は文字表示距離 Lによって延長され得 る最大の時間である。
また、 関数 f 2 ( L ) は以下の (式 7 ) によって示される。
f 2 ( L ) =l/(1+exp (- ((L-LA)/(LC-LA) -0.5)/0.1)) · ■ ■ (式 7 )
L Aは基準距離を示し、 L Cは最大距離を示す。 例えば、 基準距離は 0 c mであり、 最大距離は 1 0 c mである。
このような関数 f 2 ( L ) は、 生態系のモデルとして良く利用される シグモイ ド関数である。 即ち、 このような関数 f 2 ( L ) を用いること で、 文字表示距離しに応じた眼球運動特性に適合する移動時間 T bを導 出することができる。 また、 文字表示距離 Lは以下の (式 8 ) によって示される。
L = s q r t ( ( p X - q X ) Λ 2 + (ρ ν — q y 2 ) ■ · ■ (式 8 ) P x及び p yは、 それぞれ質問欄 1 0 7 cの文字の X座標位置及び Y 座標位置を示し、 及ぴ 9 は、 それぞれエージェン ト 1 0 7 dの X 座標位置及び Y座標位置を示す。 なお、 記号 Γ s q r t J は根を示し、 s q r t ( A ) は Aの根を示す。 また、 記号 Γ Λ」 はべき乗を示し、 (Α) Λ ( Β ) は Αの Β乗を示す。
図 8は、 関数 f 2 ( L ) 及び移動時間 T b と文字表示距離 Lとの関係 を示す図である。
関数 f 2 ( L ) によって示される値は、 図 8の ( a ) に示すように、 文字表示距離しが基準距離 L Aから最大距離 L Cに変化するに伴って增 加する。 即ち、 この値は文字表示距離 Lの増加とともに、 基準距離 L A ( O c m) 付近で緩やかに増加し、 中間距離 ( 5 c m) 付近で急激に増 加し、 最大距離 L C ( 1 O c m) 付近で再び緩やかに増加する。
したがって、 移動時間 T bは、 図 8の ( b ) に示すように、 文字表示 距離しの増加に伴って、 基準距離 L A付近で緩やかに増加し、 中間距離 付近で急激に増加し、 最大距離 L C付近で再び緩やかに増加する。
なお、 上述では文字表示距離 Lを、 エージェン ト 1 0 7 dの位置から 質問欄 1 0 7 cの文字までの距離としたが、 エージェント 1 0 7 dが表 示されていない場合には、 画面の中央を注視点として、 その中央から文 字までの距離と しても良い。
次に、 焦点合わせ诗間 T cについて詳細に説明する。
この焦点合わせ時間 T cは、 移動距離 T bと同様、 文字表示距離しに よつて定められる。
遅延部 1 0 5は、以下の(式 9 )から焦点合わせ時間 T cを導出する。
T c = t 0 + α 2 · - - (式 9 ) t 0は、 文字表示距離 Lが 0のときに要する一定の時間である。 焦点 合わせ時間 T cは、 この時間 t 0に対して、 文字表示距離しに応じて変 化する時間 2を加算することによって導出される。
遅延部 1 0 5は、 以下の (式 1 0) から時間 0ί 2を導出する。
Οί 2 = t 3 * f 2 ( L ) ■ ■ ■ (式 1 0 )
t 3は、 文字表示距離しによって延長され得る最大の時間である。 関 数 f 2 ( L ) は上述の (式 7 ) によって示される。
このような関数 f 2 ( L ) を用いることで、 文字表示距離しに応じた 眼球運動特性に適合する焦点合わせ時間 T cを導出することができる。 図 9は、 関数 f 2 ( L ) 及び焦点合わせ時間 T cと文字表示距離 Lと の関係を示す図である。
関数 f 2 ( L ) によって示される値は、 図 9の ( a ) に示すように、 文字表示距離 Lが基準距離し Aから最大距離 L Cに変化するに伴って増 加する。
したがって、 焦点合わせ時間 T cは、 図 9の ( b ) に示すように、 文 字表示距離 Lの増加に伴って、 基準距離 L A付近で緩やかに増加し、 中 間距離付近で急激に増加し、最大距離 L C付近で再び緩やかに増加する。 遅延部 1 0 5は、上述のような運動開始時間 T a と、移動時間 T bと、 焦点合わせ時間 T c とを考慮して、 以下に示す (式 1 1 ) から遅延時間 Tを導出する。
T = t O— 0 + α 1 + 2 · · · (式 1 1 )
このように、 運動開始時間 T a と移動時間 T b と焦点合わせ時間 T c とを考慮して遅延時間 Tが導出されることにより、 遅延時間 Tを人の眼 球の動きに応じた正確な時間とすることができる。
このように本実施の形態では、 ユーザが文字を視認するための動作に 要する遅延時間 Tを推定し、 文字を表示させてから遅延時間 Tだけ経過 した後に、 音声を出力するため、 ユーザは、 文字の認識を開始すると同 時に音声の認識も開始することができる。 その結果、 文字と音声による 情報をユーザに対して確実に伝えてュ一ザとの間のインターフェースの 頑健性を向上することができる。
ここで、 P D A (personal digital assistant) などの携帯端末が備 えるディスプレイのように、表示部 1 0 7の表示画面が小さい場合には、 文字表示距離 Lに関わりなく時間 α 1 と時間び 2をそれぞれ一定の時間 と しても良い。 即ち、 時間 1 及び時間 2を、 文字表示距離 Lの変化 に応じて取り得る平均的な時間とする。 このように、 時間 1 及び時間 « 2をそれぞれ平均的な時間とすることにより、 遅延部 1 0 5は遅延時 間 Τを以下の (式 1 2 ) から導出する。
Τ = t 0 — a 0 + average ( α 1 ) + average ( 2 ) ■ ■ ■ (式 1 2 ) average ( 1 ) は時間 1 の平均時間を示し、 average ( α 2 ) は時 間 α 2の平均時間を示す。
このように平均時間を用いることにより、 遅延時間 Τの導出のための パラメータ数を削減して、 計算処理を簡略化することができる。 また、 その結果、 遅延時間 τの算出速度を速めることができ、 さらに遅延部 1 0 5の構成を簡略化することができる。
また、 遅延時間 Τの上限を決めておく ことで、 遅延時間 Τが大きくな りすぎることを避ける事ができる。
(変形例 1 )
次に、 上記本実施の形態における音声出力装置の第 1 の変形例につい て説明する。
本変形例にかかる音声出力装置は、 各ユーザに応じた遅延時間を推定 するものであって、 具体的には各ユーザの年齢に応じて推定する。
一般に、 加齢とともに眼球を動かすタイミングや移動速度、 焦点を合 わせるスピ一ドは遅くなるため、 運動開始時間 T a と移動時間 T bと焦 点合わせ時間 T cも加齢とともに長くなる。 そこで本変形例にかかる音 声出力装置は、ユーザの特徴を示す年齢を考慮して遅延時間を推定する。 図 1 0は、変形例 1 にかかる音声出力装置の構成を示す構成図である。 変形例 1 にかかる音声出力装置 1 O O aは、 マイク 1 0 1 と、 音声処 理部 1 0 2と、 伝達情報生成部 1 0 3と、 計時部 1 0 4と、 遅延部 1 0 5 a と、音声出力部 1 0 6と、表示部 1 0 7 と、カードリーダ 1 0 9と、 個人情報蓄積部 1 0 8とを備える。
カードリーダ 1 0 9は、 ユーザによって音声出力装置 1 O O aに揷入 されるカード 1 0 9 aから、個人情報である年齢や生年月日を読み出し、 その読み出した個人情報を個人情報蓄積部 1 0 8に格納する。
遅延部 1 0 5 aは、 まず、 上述のように運動開始時間 T a と移動時間 T bと焦点合わせ時間 T c とを考慮した遅延時間 Tを導出する。 そして 遅延部 1 0 5 aは、 個人情報蓄積部 1 0 8に格納されている個人情報を 参照し、 遅延時間 Tからその個人情報を考慮した個人別遅延時間 T 1 を 導出する。 さらに、 遅延部 1 0 5 aは、 装置伝達情報を音声出力部 1 0 6から音声で出力させてから、 その個人別遅延時間 T 1 の経過後に、 そ の装置伝達情報を表示部 1 0 7に文字で表示させる。
遅延部 1 0 5は、 以下の (式 1 3 ) から個人別遅延時間 T 1 を導出す る。
T 1 = T + Οί 3 ■ - - (式 1 3 )
個人別遅延時間 Τ 1 は、 遅延時間 Τに対して、 年齢に応じて変化する 時間 3を加算することによって導出される。
遅延部 1 0 5は、 以下の (式 1 4 ) から時間 α 3を導出する。
0i 3 = t 4 * f 3 (M) . ' ' (式 1 4 )
Mは年齢であり、 t 4は年齢 Mによって延長され得る最大の時間であ る。
また、 関数 f 3 (M) は以下の (式 1 5 ) によって示される。
f 3 (M) =1/(1+exp (- ((M-20)/(60-20) -0.5)/0.1)) ■ ' ■ (式 1 5 )
このような関数 f 3 ( M )で示される値は加齢とともに増加するため、 その増加に伴って個人別遅延時間 T 1 も増加する。
図 1 1 は、 関数 f 3 (M) 及び個人別遅延時間 T 1 と年齢 Mとの関係 を示す図である。
関数 f 3 (M) によって示される値は、 図 1 1 の ( a ) に示すように、 年齢 Mが 2 0歳から 6 0歳に変化するに伴って増加する。 即ち、 この値 は加齢とともに、 運動能力が活発な 2 0歳 (基準年齢) 付近で緩やかに 増加し、 4 0歳 (中間年齢) 付近で急激に増加し、 運動能力が衰えた 6 0歳 (最大年齢) 付近で再び緩やかに増加する。
したがって、 個人別遅延時間 T 1 は、 図 1 1 の ( b ) に示すように、 加齢に伴って、 基準年齢付近で緩やかに増加し、 中間年齢付近で急激に 増加し、 最大年齢付近で再び緩やかに増加する。
このように本変形例では、 ユーザの年齢などを考慮して個人別遅延時 間 T 1 を導出し、 文字を表示してからその個人別遅延時間 T 1 だけ経過 した後に音声を出力するため、 各ユーザごとにインターフ: c—スの頑健 性の向上を図ることができる。
なお、 本変形例では、 個人情報と してユーザの年齢を用いたが、 ユー ザの反応速度や、 視点移動速度、 焦点合わせ速度、 俊敏性、 使用履歴な どを用いても良い。 このような場合、 上述のような反応速度などの個人 情報がカード 1 0 9 aに予め登録されており、 カードリーダ 1 0 9は、 その個人情報をカード 1 0 9 aから読み出して個人情報蓄積部 1 0 8に 格納する。 遅延部 1 0 5 aは、 個人情報蓄積部 1 0 8に格納されている 反応速度などの個人情報を参照して、 遅延時間 Tからその反応速度など を考慮した個人別遅延時間 Τ 1 を導出する。
(変形例 2 )
次に、 上記本実施の形態における音声出力装置の第 2の変形例につい て説明する。
本変形例にかかる音声出力装置は、 ユーザの慣れに応じた遅延時間を 推定するものであって、具体的にはユーザの操作回数に応じて推定する。 一般に、 ユーザが音声出力装置を操作する回数が多くなると、 ユーザ はその操作に慣れるため、 運動開始時間 T a と移動時間 T b と焦点合わ せ時間 T cは短くなる。
例えば、 切符の購入などの一連の動作において、 ユーザと音声出力装 置との対話が進むにつれて、 ユーザは、 文字の表示の位置やタイ ミング を学習する。 その結果、 ユーザは、 対話つまり操作回数が増加するごと に、 文字が表示されてからその文字に焦点を合わせるまでの動作を効率 良く行うことができる。'そこで本変形例にかかる音声出力装置は、 ユー ザの操作回数を考慮して遅延時間を推定する。
図 1 2は、変形例 2にかかる音声出力装置の構成を示す構成図である。 変形例 2にかかる音声出力装置 1 O O bは、 マイク 1 0 1 と、 音声処 理部 1 0 2と、 伝達情報生成部 1 0 3と、 計時部 1 0 4と、 遅延部 1 0 5 bと、 音声出力部 1 0 6と、 表示部 1 0 7 と、 カウンタ 1 1 0とを備 える。
カウンタ 1 1 0は、 音声処理部 1 0 2から出力されるユーザ伝達情報 を取得すると、 その取得回数、 つまリューザの音声出力装置 1 0 O bに 対する操作回数を数えて、 その操作回数を遅延部 1 0 5 bに通知する。 遅延部 1 0 5 bは、 まず、 上述のように運動開始時間 T a と移動時間 T bと焦点合わせ時間 T c とを考慮した遅延時間 Tを導出する。 そして 遅延部 1 0 5 bは、 カウンタ 1 1 0から通知される操作回数を参照し、 遅延時間 Tからその操作回数を考慮した慣れ遅延時間 T 2を導出する。 さらに、 遅延部 1 0 5 bは、 装置伝達情報を音声出力部 1 0 6から音声 で出力させてから、 その慣れ遅延時間 T 2の経過後に、 その装置伝達情 報を表示部 1 0 7に文字で表示させる。
遅延部 1 0 5 bは、 以下の (式 1 6 ) から慣れ遅延時間 T 2を導出す る。
T 2 = T - Of 4 - ■ ' (式 1 6 )
慣れ遅延時間 T 2は、 遅延時間丁から、 操作回数に応じて変化する時 間 4を減算することによって導出される。
遅延部 1 0 5 bは、 以下の (式 1 7 ) から時間 0? 4を導出する。
α 4 = t 5 * f 4 ( K ) ■ ■ ' (式 1 7 )
Κは操作回数であり、 t 5は操作回数 Kによって短縮可能な最大の時 間である。
また、 関数 f 4 ( K ) は以下の (式 1 8 ) によって示される。
f 4 ( K ) = 1/(1+exp (- (K/KG-0.5)/0.1)) ' ■ ■ (式 1 8 )
ここで K Cは、 慣れ遅延時間 T 2が最短となるような操作回数の最大 値である。
このような関数 f 4 ( K ) で示される値は操作回数 Kの増加に伴って 増加するため、慣れ遅延時間 T 2は操作回数 Kの増加に伴って減少する。 図 1 3は、 関数 f 4 ( K ) 及び慣れ遅延時間 T 2と操作回数 Kとの関 係を示す図である。
関数 f 4 ( K ) によって示される値は、 図 1 3の ( a ) に示すように、 操作回数 Kが 0回 (基準回数) から K C回 (最大回数) へ変化するに伴 つて増加する。 即ち、 この値は操作回数 Kの増加とともに、 全く不慣れ な状態の 0回付近で緩やかに増加し、 適度に慣れ始めた状態の K Cノ 2 回 (中間回数) 付近で急激に増加し、 十分に慣れた状態の K C回 (最大 回数) 付近で再び緩やかに増加する。
したがって、 慣れ遅延時間 T 2は、 図 1 3の ( b ) に示すように、 操 作回数 Kの増加に伴って、 基準回数付近で緩やかに減少し、 中間回数付 近で急激に減少し、 最大回数付近で再び緩やかに減少する。
このように本変形例では、 ユーザの慣れを考慮して慣れ遅延時間 T 2 を導出し、文字を表示してからその慣れ遅延時間 T 2だけ経過した後に、 音声を出力するため、 ユーザの慣れに適したインタ一フェースの頑健性 を保つことができる。
(変形例 3 )
次に、 上記本実施の形態における音声出力装置の第 3の変形例につい て説明する。
本変形例にかかる音声出力装置は、 変形例 2と同様、 ユーザの慣れに 応じた遅延時間を推定するものであって、 具体的にはユーザの操作時間 に応じて推定する。
一般に、 ユーザが音声出力装置を操作する時間が長くなると、 ユーザ はその操作に慣れるため、 運動開始時間 T a と移動時間 T bと焦点合わ せ時間 T cは短くなる。 そこで本変形例にかかる音声出力装置は、 ユー ザの操作時間を考慮して遅延時間を推定する。
変形例 3にかかる音声出力装置は、 図 1 2に示す変形例 2の音声出力 装置 1 O O bと同様の構成を有するが、 遅延部 1 0 5 bとカウンタ 1 1 0の動作が異なる。
本変形例にかかるカウンタ 1 1 0は、 タイムカウンタとしての機能を 有し、 音声出力装置 1 0 0 b とユーザとの間で対話が開始された後に音 声処理部 1 0 2から最初のユーザ伝達情報を取得すると、 その取得した ときからの経過時間、 つまり操作時間を計測する。 そしてカウンタ 1 1 0はその操作時間を遅延部 1 0 5 bに通知する。
遅延部 1 05 bは、 まず、 上述のように運動開始時間 T a と移動時間 T bと焦点合わせ時間 T c とを考慮した遅延時間 Tを導出する。 そして 遅延部 1 0 5 bは、 カウンタ 1 1 0から通知される操作時間を参照し、 遅延時間 Tからその操作時間を考慮した慣れ遅延時間 T 3を導出する。 さらに、 遅延部 1 0 5 bは、 装置伝達情報を音声出力部 1 0 6から音声 で出力させてから、 その慣れ遅延時間 T 3の経過後に、 その装置伝達情 報を表示部 1 0 7に文字で表示させる。
遅延部 1 0 5 bは、 以下の (式 1 9 ) から慣れ遅延時間 T 3を導出す る。
T 3 = T - α 5 ■ - - (式 1 9 )
慣れ遅延時間 T 3は、 遅延時間丁から、 操作時間に応じて変化する時 間 5を減算することによって導出される。
遅延部 1 0 5 bは、 以下の (式 20 ) から時間 α 5を導出する。
Qf 5 = t 6 * f 5 ( P ) ■ ■ ■ (式 2 0)
Pは操作時間であり、 t 6は操作時間 Pによって短縮可能な最大の時 間である。
また、 関数 f 5 ( P ) は以下の (式 2 1 ) によって示される。
f 5 ( P ) =1/(1+exp (- (P/PG-0.5)/0.1)) - ■ . (式 2 1 )
ここで P Cは、 慣れ遅延時間 T 3が最短となるような操作時間 Pの最 大値である。
このような関数 f 5 ( P ) で示される値は操作時間 Pの増加に伴って 増加するため、慣れ遅延時間 T 3は操作時間 Pの増加に伴って減少する。
このように本変形例では、 変形例 2と同様、 ユーザの慣れを考慮して 慣れ遅延時間 T 3を導出し、 文字を表示してからその慣れ遅延時間 T 3 だけ経過した後に音声を出力するため、 ユーザの慣れに適したインター フェースの頑健性を保つことができる。
なお、 本変形例では、 音声処理部 1 0 2がユーザ伝達情報を出力した タイミング、 即ちユーザが音声を発したタィミングで操作時間 Pの計測 を開始したが、 音声出力装置 1 0 0 bに電源が投入されたタィミング、 又は、 開始ボタン 1 0 7 f が選択されたタイミングで操作時間 Pの計測 を開始しても良い。
(変形例 4 )
次に、 本実施の形態における運動開始時間 T aの導出方法に関する第 4の変形例について説明する。
一般に、 文字のサイズだけではなく、 文字の表示位置によっても運動 開始時間 T aは変化する。 つまり、 表示される文字の位置が、 ユーザの 注視点に近ければ近いほど、 ユーザはその文字に早く気づくため、 運動 開始時間 T aは短くなる。
本変形例にかかる遅延部 1 0 5は、 文字表示距離しに基づいて運動開 始時間 T aを以下の (式 2 2 ) から導出する。
T a = t 0 + 6 - ' - (式 2 2 )
t 0は、文字表示距離 Lが 0のときに要する一定の時間である。即ち、 運動開始時間 T aは、 この時間 t 0に対して、 文字表示距離しによって 変化する時間 6を加算することによって導出される。
遅延部 1 0 5は、 以下の (式 2 3 ) から 6を導出する。
α 6 = t 7 * f 2 ( L ) , ■ ■ (式 2 3 )
t 7は文字表示距離しによって延長され得る最大の時間である。また、 関数 f 2 ( L ) は (式 7 ) によって示される。
このように本変形例では、 文字表示距離 Lを考慮した運動開始時間 T aに基づく遅延時間 Tを導出し、 文字を表示してからその遅延時間 Tだ け経過した後に音声を出力するため、 各文字表示距離 Lに適したインタ —フェースの頑健性を保つことができる。
(変形例 5 )
次に、 本実施の形態における運動開始時間 T aの導出方法に関する第 5の変形例について説明する。
—般に、 ユーザの注視点と、 表示される文字の色のコン トラス トとが 大きく異なれば、 それだけユーザはその文字に早く気づ〈ため、 運動開 始時間 T aは短くなる。
本変形例にかかる遅延部 1 0 5は、 注視点と文字のコン トラス トに基 づいて運動開始時間 T aを以下の (式 2 4 ) から導出する。
T a = t 0 - α 7 - - - (式 2 4 )
t 0は、 コン トラス トを限りなく小さく したときに要する一定の時間 である。 即ち、 運動開始時間 T aは、 この時間 t 0から、 コ ン トラス ト によって変化する時間《 7を減算することによって導出される。
遅延部 1 0 5は、 以下の (式 2 5 ) から 7を導出する。
Qi 7 = t 8 * f 6 (Q ) - ' - (式 2 5 )
Qはコントラス トを示し、 t 8はコン トラス ト Qによって短縮可能な 最大の時間である。
また、 関数 f 6 ( Q ) は以下の (式 2 6 ) によって示される。
f 6 ( Q) =1/(l+exp (- ((Q-QA)/(QC-QA) -0.5)/0.1)) ' ■ · (式 2 6 )
Q Aは運動開始時間 T a を決めるための基準のコン トラス トであり、 Q Cは運動開始時間 T aを決めるための最大のコン トラス トである。 このように本変形例では、 コン トラス トを考慮した運動開始時間 T a に基づ〈遅延時間 Tを導出し、 文字を表示してからその遅延時間 Tだけ 経過した後に音声を出力するため、 各コントラス トに適したインターフ エースの頑健性を保つことができる。 (変形例 6 )
次に、 本実施の形態における運動開始.時間 T aの導出方法に関する第 6の変形例について説明する。
一般に、 文字を赤で表示したり、 点滅させたりすることにより、 ユー ザはその文字に早く気づくため、 運動開始時間 T aは短くなる。
本変形例にかかる遅延部 1 0 5は、 文字の表示態様の強調度合いに基 づいて運動開始時間 T a を以下の (式 2 7 ) から導出する。
T a = t 0 - α 8 - - - (式 2 7 )
t 0は、 表示形態の強調度合いを限リなく小さく したときに要する一 定の時間である。 即ち、 運動開始時間 T aは、 この時間 t Oから、 強調 度合いによって変化する時間 8を減算することによって導出される。 遅延部 1 0 5は、 以下の (式 2 8 ) から QT 8を導出する。
α 8 = t 9 * f 7 ( R ) . ' · (式 2 8 )
Rは強調度合いを示し、 t 9は強調度合い Rによって短縮可能な最大 の時間である。
また、 関数 f 7 ( R) は以下の (式 2 9 ) によって示される。
f 7 ( R ) = 1/(1+exp (- ((R-RA)/(RC-RA) -0.5)/0.1) ) ■ · ■ (式 2 9 )
R Aは運動開始時間 T aを決めるための基準の強調度合いであリ、 R Cは運動開始時間 T aを決めるための最大の強調度合いである。
このように本変形例では、 文字の強調度合いを考慮した運動開始時間 T aに基づく遅延時間 Tを導出し、 文字を表示してからその遅延時間 T だけ経過した後に音声を出力するため、 文字の各強調度合いに適したィ ンタ一フエ一スの頑健性を保つことができる。
以上、 本発明について実施の形態及び変形例を用いて説明したが、 本 発明はこれらに限定されるものではない。 例えば、 実施の形態及び変形例では、 遅延時間 Tを導出するために、 運動開始時間 T a と移動時間 T b と焦点合わせ時間 T c とを全て考慮し たが、 これらのうちの少なく とも 1 つを考慮して遅延時間丁を導出して も良い。
また、変形例 1 では、ユーザの個人情報に基づいて遅延時間を導出し、 変形例 2では、 ユーザの慣れに基づいて遅延時間を導出したが、 ユーザ の個人情報及び慣れの双方に基づいて遅延時間を導出しても良い。 また、 実施の形態及び変形例では、 音声出力装置を音声の出力及び文 字の表示によリ切符の販売を行う装置と して説明したが、 音声の出力及 び文字の表示を行うものであれば他の動作を行う装置であっても良い。 例えば、 テレビゃ、 カーナピゲーシヨンシステムの端末、 携帯電話、 携 末、 パ ―ソナルコンビユ ータ、 電話機、 ファックス、 電子レンジ、 冷蔵庫、 掃除機、 電子辞書 、 電子翻訳機などと して音声出力装置を構成 しても良い。 産業上の利用の可能性
本発明に係る音声出力装置は、 文字と音声による情報をユーザに対し て確実に伝えてユーザとの間のインターフェースの頑健性を向上するこ とができ、 例えばユーザの音声に対して音声及び文字で応答することに より切符などを販売する音声応答装置などに有用である。

Claims

1 . ユーザに対して伝達すべき伝達情報を文字で表示する文字表示手 段と、
前記文字表示手段に文字が表示されてから、 ユーザが前記文字を視認 するための動作に要する遅延時間が経過したときに、 前記伝達情報を音 請
声で出力する音声出力手段と
を備えることを特徴とする音声出力装置。
の 範
2 . 前記音声出力装置は、 さらに、
前記文字表示手段に表示される文字の表示態様に応じて前記遅延時間 を推定する遅延推定手段を備え、
前記音声出力手段は、
前記文字表示手段に文字が表示されてから、 前記遅延推定手段により 推定された遅延時間が経過したときに、 前記伝達情報を音声で出力する ことを特徴とする請求の範囲第 1項記載の音声出力装置。
3 . 前記遅延推定手段は、
前記文字表示手段により文字が表示されてから、 ユーザの視点が前記 文字に移動を開始するまでの開始時間を含むように前記遅延時間を推定 する
ことを特徴とする請求の範囲第 2項記載の音声出力装置。
4 . 前記遅延推定手段は、 さら
ユーザの視点が移動を開始してから前記文字に到達するまでの移動時 間を含むように前記遅延時間を推定する ことを特徴とする請求の範囲第 3記載の音声出力装置。
5 . 前記遅延推定手段は、 さらに、
ユーザの視点が前記文字に到達してから前記文字に焦点が合うまでの 焦点合わせ時間を含むように前記遅延時間を推定する
ことを特徴とする請求の範囲第 4項記載の音声出力装置。
6 . 前記音声出力装置は、 さらに、
ユーザの特徴を示す個人情報を取得する個人情報取得手段を備え、 前記遅延推定手段は、
前記個人情報取得手段により取得された個人情報に基づいて、 前記ュ 一ザに応じた遅延時間を推定する
ことを特徴とする請求の範囲第 5項記載の音声出力装置。
7 . 前記個人情報取得手段は、 前記ユーザの年齢を前記個人情報とし て取得し、
前記遅延推定手段は、
前記個人情報取得手段によリ取得された年齢に基づいて、 前記ユーザ に応じた遅延時間を推定する
ことを特徴とする請求の範囲第 6項記載の音声出力装置。
8 . 前記個人情報取得手段は、 前記ユーザの眼球を動かす眼球速度を 前記個人情報と して取得し、
前記遅延推定手段は、
前記個人情報取得手段により取得された眼球速度に基づいて、 前記ュ 一ザに応じた遅延時間を推定する ことを特徴とする請求の範囲第 6項記載の音声出力装置。
9 . 前記音声出力装置は、 さらに、
ユーザによる操作に応じて、 前記文字表示手段に文字を表示させると ともに、 前記音声出力手段に音声を出力させる操作手段と、
ユーザの前記操作手段に対する操作の慣れ度合いを特定する慣れ特定 手段とを備え、
前記遅延推定手段は、
前記慣れ特定手段により特定された慣れ度合いに基づいて、 前記ユー ザの慣れに応じた遅延時間を推定する
ことを特徴とする請求の範囲第 5項記載の音声出力装置。
1 0 . 前記慣れ特定手段は、 前記操作手段に対するユーザの操作回 を前記慣れ度合いとして特定する
ことを特徴とする請求の範囲第 9項記載の音声出力装置。
1 1 . 前記慣れ特定手段は、 前記操作手段に対するユーザの操作時間 を前記慣れ度合いとして特定する
ことを特徴とする請求の範囲第 9項記載の音声出力装置。
1 2 . 前記遅延推定手段は、
ユーザの注意を引き付ける前記音声出力装置の注視点から、 前記文字 表示手段により表示された文字までの文字表示距離に基づいて、 前記焦 点合わせ時間を特定する
ことを特徴とする請求の範囲第 5項記載の音声出力装置。
1 3 . 前記遅延推定手段は、
前記注視点と前記文字のそれぞれの位置から前記文字表示距離を導出 し、 前記文字表示距離に基づいて前記焦点合わせ時間を特定する ことを特徴とする請求の範囲第 1 2項記載の音声出力装置。
1 4 . 前記文字表示手段は、
エージェン トを前記注視点として表示し、
前記遅延推定手段は、 前記エージェン トを起点として文字表示距離を 導出する
ことを特徴とする請求の範囲第 1 2項記載の音声出力装置。
1 5 . 前記遅延推定手段は、
シグモイ ド関数を利用して前記遅延時間を推定する
ことを特徴とする請求の範囲第 5項記載の音声出力装置。
1 6 . 前記遅延推定手段は、
ユーザの注意を引き付ける前記音声出力装置の注視点から、 前記文字 表示手段により表示された文字までの文字表示距離に基づいて、 前記移 動時間を特定する
ことを特徴とする請求の範囲第 4項記載の音声出力装置。
1 7 . 前記遅延推定手段は、
前記文字表示手段により表示される文字のサイズに基づいて、 前記開 始時間を特定する
ことを特徴とする請求の範囲第 3項記載の音声出力装置。
1 8 . 前記遅延推定手段は、
ユーザの注意を引き付ける前記音声出力装置の注視点から、 前記文字 表示手段により表示された文字までの文字表示距離に基づいて、 前記開 始時間を特定する
ことを特徴とする請求の範囲第 3項記載の音声出力装置。
1 9 . 前記遅延推定手段は、
ユーザの注意を引き付ける前記音声出力装置の注視点と、 前記文字表 示手段により表示された文字とのコン トラス 卜に基づいて、 前記開始時 間を特定する
ことを特徴とする請求の範囲第 3項記載の音声出力装置。
2 0 . 前記文字表示手段は、 文字を点滅させて表示し、
前記遅延推定手段は、
前記文字表示手段により表示された文字の点滅度合いに基づいて、 前 記開始時間を特定する
ことを特徴とする請求の範囲第 3項記載の音声出力装置。
2 1 . 前記遅延推定手段は、
ユーザの視点が移動を開始してから前記文字に到達するまでの移動時 間を含むように、 前記遅延時間を推定する
ことを特徴とする請求の範囲第 2項記載の音声出力装置。
2 2 . 前記遅延推定手段は、
ユーザの視点が前記文字に到達してから前記文字に焦点が合うまでの 焦点合わせ時間を含むように前記遅延時間を推定する ことを特徴とする請求の範囲第 2項記載の音声出力装置
2 3 . 情報処理装置が音声を出力する方法であって、
人に対して伝達すべき伝達情報を文字で表示する文字表示ステップと、 前記文字表示ステップで文字が表示されてから、 人が前記文字を視認 するための動作に要する遅延時間が経過したときに、 前記伝達情報を音 声で出力する音声出力ステップと
を含むことを特徴とする音声出力方法。
2 4 . 前記音声出力方法は、 さらに、
前記文字表示ステップで表示される文字の表示態様に応じて前記遅延 時間を推定する遅延推定ステップを含み、
前記音声出カステツプでは、
前記文字表示ステップで文字が表示されてから、 前記遅延推定ステツ プで推定された遅延時間が経過したときに、 前記伝達情報を音声で出力 する
ことを特徴とする請求の範囲第 2 3記載の音声出力方法。
2 5 . 前記遅延推定ステップでは、
前記文字表示ステップで文字が表示されてから、 前記文字にユーザの 視点が移動を開始するまでの開始時間を含むように前記遅延時間を推定 する
ことを特徴とする請求の範囲第 2 4項記載の音声出力方法。
2 6 . 前記遅延推定ステップでは、 さらに、
ユーザの視点が移動を開始してから前記文字に到達するまでの移動時 間を含むように前記遅延時間を推定する
ことを特徴とする請求の範囲第 2 5項記載の音声出力方法。
2 7 . 前記遅延推定ステップでは、 さらに、
ユーザの視点が前記文字に到達してから前記文字に焦点が合うまでの 焦点合わせ時間を含むように前記遅延時間を推定する
ことを特徴とする請求の範囲第 2 6項記載の音声出力方法。
2 8 . 人に対して伝達すべき伝達情報を文字で表示する文字表示ステ ップと、
前記文字表示ステップで文字が表示されてから、 人が前記文字を視認 するための動作に要する遅延時間が経過したときに、 前記伝達情報を音 声で出力する音声出力ステップと
をコンピュータに実行させることを特徴とするプログラム。
PCT/JP2004/006065 2003-05-21 2004-04-27 音声出力装置及び音声出力方法 Ceased WO2004104986A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US10/542,947 US7809573B2 (en) 2003-05-21 2004-04-27 Voice output apparatus and voice output method
JP2005504492A JP3712207B2 (ja) 2003-05-21 2004-04-27 音声出力装置及び音声出力方法

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2003-143043 2003-05-21
JP2003143043 2003-05-21

Publications (1)

Publication Number Publication Date
WO2004104986A1 true WO2004104986A1 (ja) 2004-12-02

Family

ID=33475114

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2004/006065 Ceased WO2004104986A1 (ja) 2003-05-21 2004-04-27 音声出力装置及び音声出力方法

Country Status (4)

Country Link
US (1) US7809573B2 (ja)
JP (1) JP3712207B2 (ja)
CN (1) CN100583236C (ja)
WO (1) WO2004104986A1 (ja)

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010214572A (ja) * 2009-03-19 2010-09-30 Denso Wave Inc ロボット制御命令入力装置
JP2017062611A (ja) * 2015-09-24 2017-03-30 富士ゼロックス株式会社 情報処理装置及びプログラム
JPWO2021085242A1 (ja) * 2019-10-30 2021-05-06

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8219067B1 (en) * 2008-09-19 2012-07-10 Sprint Communications Company L.P. Delayed display of message
US8456420B2 (en) * 2008-12-31 2013-06-04 Intel Corporation Audible list traversal
CN101807427A (zh) * 2010-02-26 2010-08-18 中山大学 一种音频输出系统
CN101776987A (zh) * 2010-02-26 2010-07-14 中山大学 一种输出植物音频的方法
JP6548524B2 (ja) * 2015-08-31 2019-07-24 キヤノン株式会社 情報処理装置、情報処理システム、情報処理方法、及びプログラム
JP2017123564A (ja) * 2016-01-07 2017-07-13 ソニー株式会社 制御装置、表示装置、方法及びプログラム
CN105550679B (zh) * 2016-02-29 2019-02-15 深圳前海勇艺达机器人有限公司 一种机器人循环监听录音的判断方法
WO2022054407A1 (ja) * 2020-09-08 2022-03-17 パナソニック インテレクチュアル プロパティ コーポレーション オブ アメリカ 行動推定装置、行動推定方法、及び、プログラム

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06321030A (ja) * 1993-05-11 1994-11-22 Yazaki Corp 車両用音声出力装置
JPH0851891A (ja) * 1994-08-12 1996-02-27 Geochto:Kk 蚕の無菌飼育方法
JP2003099092A (ja) * 2001-09-21 2003-04-04 Chuo Joho Kaihatsu Kk 携帯型情報端末を利用した音声入力システム、並びに、これを使用した棚卸又は受発注用の商品管理システム、点検管理システム、訪問介護管理システム、及び看護管理システム
JP2003108171A (ja) * 2001-09-27 2003-04-11 Clarion Co Ltd 文書読み上げ装置

Family Cites Families (18)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4884972A (en) * 1986-11-26 1989-12-05 Bright Star Technology, Inc. Speech synchronized animation
US5303327A (en) * 1991-07-02 1994-04-12 Duke University Communication test system
JPH05216618A (ja) 1991-11-18 1993-08-27 Toshiba Corp 音声対話システム
JP3667615B2 (ja) 1991-11-18 2005-07-06 株式会社東芝 音声対話方法及びそのシステム
JPH0792993A (ja) * 1993-09-20 1995-04-07 Fujitsu Ltd 音声認識装置
JPH0854894A (ja) 1994-08-10 1996-02-27 Fujitsu Ten Ltd 音声処理装置
IL110883A (en) * 1994-09-05 1997-03-18 Ofer Bergman Reading tutorial system
US6073103A (en) * 1996-04-25 2000-06-06 International Business Machines Corporation Display accessory for a record playback system
US5850211A (en) * 1996-06-26 1998-12-15 Sun Microsystems, Inc. Eyetrack-driven scrolling
JP3804188B2 (ja) 1997-06-09 2006-08-02 ブラザー工業株式会社 文章読み上げ装置
JP3861413B2 (ja) 1997-11-05 2006-12-20 ソニー株式会社 情報配信システム、情報処理端末装置、携帯端末装置
JP3125746B2 (ja) 1998-05-27 2001-01-22 日本電気株式会社 人物像対話装置及び人物像対話プログラムを記録した記録媒体
GB2353927B (en) * 1999-09-06 2004-02-11 Nokia Mobile Phones Ltd User interface for text to speech conversion
US20040006473A1 (en) * 2002-07-02 2004-01-08 Sbc Technology Resources, Inc. Method and system for automated categorization of statements
JP2003241779A (ja) 2002-02-18 2003-08-29 Sanyo Electric Co Ltd 情報再生装置および情報再生方法
US6800274B2 (en) * 2002-09-17 2004-10-05 The C.P. Hall Company Photostabilizers, UV absorbers, and methods of photostabilizing a sunscreen composition
US7151435B2 (en) * 2002-11-26 2006-12-19 Ge Medical Systems Information Technologies, Inc. Method and apparatus for identifying a patient
US20040190687A1 (en) * 2003-03-26 2004-09-30 Aurilab, Llc Speech recognition assistant for human call center operator

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH06321030A (ja) * 1993-05-11 1994-11-22 Yazaki Corp 車両用音声出力装置
JPH0851891A (ja) * 1994-08-12 1996-02-27 Geochto:Kk 蚕の無菌飼育方法
JP2003099092A (ja) * 2001-09-21 2003-04-04 Chuo Joho Kaihatsu Kk 携帯型情報端末を利用した音声入力システム、並びに、これを使用した棚卸又は受発注用の商品管理システム、点検管理システム、訪問介護管理システム、及び看護管理システム
JP2003108171A (ja) * 2001-09-27 2003-04-11 Clarion Co Ltd 文書読み上げ装置

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
TAMURA H.: "Human interface", OHMSHA LTD., 30 May 1998 (1998-05-30), pages 23-33,69 - 79,212-213,383-411, XP002983541 *

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010214572A (ja) * 2009-03-19 2010-09-30 Denso Wave Inc ロボット制御命令入力装置
JP2017062611A (ja) * 2015-09-24 2017-03-30 富士ゼロックス株式会社 情報処理装置及びプログラム
JPWO2021085242A1 (ja) * 2019-10-30 2021-05-06
WO2021085242A1 (ja) * 2019-10-30 2021-05-06 ソニー株式会社 情報処理装置、及びコマンド処理方法
US20220357915A1 (en) * 2019-10-30 2022-11-10 Sony Group Corporation Information processing apparatus and command processing method
JP7533472B2 (ja) 2019-10-30 2024-08-14 ソニーグループ株式会社 情報処理装置、及びコマンド処理方法
US12182475B2 (en) 2019-10-30 2024-12-31 Sony Group Corporation Information processing apparatus and command processing method

Also Published As

Publication number Publication date
CN100583236C (zh) 2010-01-20
US20060085195A1 (en) 2006-04-20
JPWO2004104986A1 (ja) 2006-07-20
US7809573B2 (en) 2010-10-05
JP3712207B2 (ja) 2005-11-02
CN1759436A (zh) 2006-04-12

Similar Documents

Publication Publication Date Title
JP3514372B2 (ja) マルチモーダル対話装置
US20230306968A1 (en) Digital assistant for providing real-time social intelligence
EP3217254A1 (en) Electronic device and operation method thereof
US20120226503A1 (en) Information processing apparatus and method
KR20180111197A (ko) 정보 제공 방법 및 이를 지원하는 전자 장치
JP6407521B2 (ja) 診療支援装置
WO2007041223A2 (en) Automated dialogue interface
US11244682B2 (en) Information processing device and information processing method
US20200005332A1 (en) Systems, devices, and methods for providing supply chain and ethical sourcing information on a product
CN114581955B (zh) 近视防控方法、装置、系统、存储介质和设备
WO2004104986A1 (ja) 音声出力装置及び音声出力方法
CN111641677A (zh) 消息提醒方法、消息提醒装置及电子设备
JP7021488B2 (ja) 情報処理装置、及びプログラム
JP2020160641A (ja) 仮想人物選定装置、仮想人物選定システム及びプログラム
CN112312211A (zh) 提示方法和装置
JP2017117184A (ja) ロボット、質問提示方法、及びプログラム
EP4068119A1 (en) Model training method and apparatus for information recommendation, electronic device and medium
KR20220023211A (ko) 대화 텍스트에 대한 요약 정보를 생성하는 전자 장치 및 그 동작 방법
KR20190067433A (ko) 텍스트-리딩 기반의 리워드형 광고 서비스 제공 방법 및 이를 수행하기 위한 사용자 단말
CN110634570A (zh) 一种诊断仿真方法及相关装置
JP2001350904A (ja) 電子機器及び営業用電子機器システム
JP7582355B2 (ja) 情報処理装置、制御方法、及びプログラム
CN110751951A (zh) 基于智能镜子的握手交互方法及系统、存储介质
KR102324379B1 (ko) 휘트니스 상품 거래 플랫폼을 운용하는 서버 및 그 동작 방법
CN115171284A (zh) 一种老年人关怀方法及装置

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: 2005504492

Country of ref document: JP

AK Designated states

Kind code of ref document: A1

Designated state(s): AE AG AL AM AT AU AZ BA BB BG BR BW BY BZ CA CH CN CO CR CU CZ DE DK DM DZ EC EE EG ES FI GB GD GE GH GM HR HU ID IL IN IS JP KE KG KP KR KZ LC LK LR LS LT LU LV MA MD MG MK MN MW MX MZ NA NI NO NZ OM PG PH PL PT RO RU SC SD SE SG SK SL SY TJ TM TN TR TT TZ UA UG US UZ VC VN YU ZA ZM ZW

AL Designated countries for regional patents

Kind code of ref document: A1

Designated state(s): BW GH GM KE LS MW MZ NA SD SL SZ TZ UG ZM ZW AM AZ BY KG KZ MD RU TJ TM AT BE BG CH CY CZ DE DK EE ES FI FR GB GR HU IE IT LU MC NL PL PT RO SE SI SK TR BF BJ CF CG CI CM GA GN GQ GW ML MR NE SN TD TG

121 Ep: the epo has been informed by wipo that ep was designated in this application
ENP Entry into the national phase

Ref document number: 2006085195

Country of ref document: US

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 10542947

Country of ref document: US

WWE Wipo information: entry into national phase

Ref document number: 20048062318

Country of ref document: CN

WWP Wipo information: published in national office

Ref document number: 10542947

Country of ref document: US

122 Ep: pct application non-entry in european phase