WO2024252668A1 - 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム - Google Patents

個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム Download PDF

Info

Publication number
WO2024252668A1
WO2024252668A1 PCT/JP2023/021547 JP2023021547W WO2024252668A1 WO 2024252668 A1 WO2024252668 A1 WO 2024252668A1 JP 2023021547 W JP2023021547 W JP 2023021547W WO 2024252668 A1 WO2024252668 A1 WO 2024252668A1
Authority
WO
WIPO (PCT)
Prior art keywords
user
virtual environment
personal
synthesis device
personal virtual
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2023/021547
Other languages
English (en)
French (fr)
Inventor
秀明 岩本
達明 伊藤
直紀 萩山
俊一 瀬古
尚司 松川
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
NTT Inc
Original Assignee
Nippon Telegraph and Telephone Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Telegraph and Telephone Corp filed Critical Nippon Telegraph and Telephone Corp
Priority to PCT/JP2023/021547 priority Critical patent/WO2024252668A1/ja
Priority to JP2025525912A priority patent/JPWO2024252668A1/ja
Publication of WO2024252668A1 publication Critical patent/WO2024252668A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R1/00Details of transducers, loudspeakers or microphones
    • H04R1/10Earpieces; Attachments therefor ; Earphones; Monophonic headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04RLOUDSPEAKERS, MICROPHONES, GRAMOPHONE PICK-UPS OR LIKE ACOUSTIC ELECTROMECHANICAL TRANSDUCERS; ELECTRIC HEARING AIDS; PUBLIC ADDRESS SYSTEMS
    • H04R3/00Circuits for transducers

Definitions

  • the embodiments relate to a personal virtual environment synthesis device, a personal virtual environment synthesis method, and a program.
  • the facial position, posture, and line of sight of the conversation participants are estimated using images obtained from an observation device.
  • the presence or absence of speech and the direction of arrival are estimated using sounds obtained from the observation device.
  • the degree of gaze at the virtual camera, the azimuth angle of the virtual camera relative to the origin of the virtual space, and control parameters that control the viewpoint of the virtual camera are calculated.
  • images of the conversation participants are projected onto a partial plane, and the projected partial plane is placed on a horizontal plane in the virtual space so as to correspond to the position of the actual conversation participants.
  • a virtual space image from the viewpoint of the virtual camera is generated using the control parameters.
  • Non-Patent Document 1 it is believed that significant noise related to the content of the work affects the work, with regard to the relationship between sound and the efficiency of intellectual work. It is also believed that being aware that conversations taking place in the same room also relate to oneself may be a factor that interferes with concentration.
  • the speaker's voice and speaker's position are estimated from data from a microphone array and an omnidirectional camera.
  • the individual the speaker faces is estimated from, for example, the speaker's face and gaze direction in the omnidirectional camera data.
  • the individual each individual faces is estimated from, for example, the speaker's face and gaze direction in the omnidirectional camera data.
  • the present invention was made with the above in mind, and its purpose is to provide a means to prevent excessive fatigue accumulation in individuals by naturally increasing concentration and relieving tension without interfering with work.
  • the personal virtual environment synthesis device of one embodiment includes a concentration level acquisition unit that calculates the concentration level of a user, and a control unit that controls the occlusion rate of the user's surrounding environment based on the calculated concentration level.
  • FIG. 1 is a diagram illustrating an example of a personal virtual environment synthesis system according to an embodiment.
  • FIG. 2 is a block diagram illustrating an example of a hardware configuration of a terminal according to the embodiment.
  • FIG. 3 is a block diagram showing an example of a hardware configuration of the personal virtual environment synthesis device according to the embodiment.
  • FIG. 4 is a block diagram showing an example of a functional configuration of the personal virtual environment synthesis device according to the embodiment.
  • FIG. 5 is a diagram illustrating an example of individual characteristic information according to the embodiment.
  • FIG. 6 is a flowchart showing an example of a personal virtual environment synthesis process in the personal virtual environment synthesis device according to the embodiment.
  • FIG. 7 is a diagram for explaining the time at which work is interrupted, which is detected in the personal virtual environment synthesis apparatus according to the embodiment.
  • FIG. 8 is a graph for explaining an example of the relationship between the concentration level and the control ratio in personal virtual environment synthesis according to the embodiment.
  • FIG. 1 is a diagram showing an example of a personal virtual environment synthesis system according to an embodiment.
  • the personal virtual environment synthesis system 1 is a system that synthesizes the virtual environments of multiple users who work within the same space A, for example.
  • the personal virtual environment synthesis system 1 includes terminals 11-1, 11-2, 11-3, 11-4, 11-5, and 11-6, a camera 12, a microphone array 13, headphones 14-1, 14-2, 14-3, 14-4, 14-5, and 14-6, and a personal virtual environment synthesis device 15.
  • terminals 11-1, 11-2, 11-3, 11-4, 11-5, and 11-6 each of terminals 11-1, 11-2, 11-3, 11-4, 11-5, and 11-6 will simply be referred to as terminal 11.
  • each of the headphones 14-1, 14-2, 14-3, 14-4, 14-5, and 14-6 will simply be referred to as the headphones 14.
  • Each terminal 11 is a personal computer or a smartphone.
  • the terminal 11 and the personal virtual environment synthesis device 15 are configured to be able to communicate with each other, for example, via a network NW.
  • Terminals 11-1, 11-2, 11-3, 11-4, 11-5, and 11-6 are used for the business of users U1, U2, U3, U4, U5, and U6, respectively.
  • Users U1, U2, U3, U4, U5, and U6 perform business within the same space A.
  • each of users U1, U2, U3, U4, U5, and U6 will simply be referred to as user U.
  • the camera 12 for example, acquires photographic data of the user U.
  • the camera 12 is, for example, an omnidirectional camera.
  • the camera 12 is configured to be able to transmit the acquired photographic data to the personal virtual environment synthesis device 15, for example, via the network NW.
  • the microphone array 13 includes multiple microphones MP.
  • the microphone array 13 acquires audio data within space A using the multiple microphones MP.
  • the audio data within space A includes, for example, audio data of conversations between users U1 to U6 (conversation sound data) and audio data of background sounds (background sound data).
  • the conversation sound data may include audio data of multiple conversations between users U1 to U6.
  • the background sound data includes, for example, data such as the operating sounds of air conditioning and printers.
  • the microphone array 13 transmits the acquired audio data to the personal virtual environment synthesis device 15, for example, via the network NW.
  • Headphones 14-1, 14-2, 14-3, 14-4, 14-5, and 14-6 are used by users U1, U2, U3, U4, U5, and U6, respectively.
  • Audio data for user U using each headphone 14 is transmitted to each headphone 14 from the personal virtual environment synthesizer 15 via, for example, the terminal 11 used by that user U. Note that audio data may also be transmitted directly from the personal virtual environment synthesizer 15 to each headphone 14.
  • the personal virtual environment synthesis device 15 is, for example, a server that manages the image data of the camera 12, the audio data acquired by the microphone array 13, and the audio data output to the headphones 14.
  • the personal virtual environment synthesis device 15 is configured to be able to separate and extract, for example, conversation sound data and background sound data based on the image data of the camera 12 and the audio data acquired by the microphone array 13.
  • the personal virtual environment synthesis device 15 also calculates the concentration level (percentage) of each user U based on, for example, the image acquired from the camera 12.
  • the personal virtual environment synthesis device 15 synthesizes audio data for the user U (audio data of the virtual environment sound) using each of the audio data separated and extracted as described above and audio data of notification to the terminal 11 (notification sound data).
  • the notification sound data includes, for example, data such as a ringtone for an email to be notified to the terminal 11.
  • the personal virtual environment synthesis device 15 then outputs the audio data of the virtual environment sound to the headphones 14 of the user U.
  • the personal virtual environment synthesis device 15 controls the size of the working window (active window) on the terminal 11 used by each user U, for example, based on the concentration level of the user U.
  • the personal virtual environment synthesis device 15 may be located inside or outside space A, and there is no particular limitation on the location of the personal virtual environment synthesis device 15.
  • FIG. 2 is a block diagram showing an example of the hardware configuration of a terminal according to an embodiment.
  • each terminal 11 includes a control circuit 111, a communication module 112, and a user interface 113.
  • Terminals 11-1, 11-2, 11-3, 11-4, 11-5, and 11-6 have the same configuration.
  • the control circuit 111 is a circuit that provides overall control of each component of the terminal 11.
  • the control circuit 111 includes a CPU (Central Processing Unit), RAM (Random Access Memory), and ROM (Read Only Memory).
  • the ROM of the control circuit 111 stores programs and the like used in various processes in the terminal 11.
  • the CPU of the control circuit 111 controls the entire terminal 11 in accordance with the programs stored in the ROM of the control circuit 111.
  • the RAM of the control circuit 111 is used as a working area for the CPU of the control circuit 111.
  • the communication module 112 is a circuit used to send and receive data between the terminal 11 and the personal virtual environment synthesis device 15.
  • the user interface 113 is an interface that handles communication between the user U and the control circuit 111.
  • the user interface 113 includes input devices and output devices.
  • the input devices include, for example, an audio microphone, a touch panel, and operation buttons.
  • the output devices include, for example, a speaker and a display.
  • each terminal 11 may include a camera (built-in camera).
  • FIG. 3 is a block diagram showing an example of the hardware configuration of the personal virtual environment synthesis device according to the embodiment.
  • the personal virtual environment synthesis device 15 includes a control circuit 151, a communication module 152, a storage 153, a drive 154, and a storage medium 155.
  • the control circuit 151 is a circuit that provides overall control of each component of the personal virtual environment synthesis device 15.
  • the control circuit 151 includes a CPU, RAM, ROM, etc.
  • the ROM of the control circuit 151 stores programs and the like used in various processes in the personal virtual environment synthesis device 15.
  • the CPU of the control circuit 151 controls the entire personal virtual environment synthesis device 15 in accordance with the programs stored in the ROM of the control circuit 151.
  • the RAM of the control circuit 151 is used as a working area for the CPU of the control circuit 151.
  • the communication module 152 is a circuit used to transmit and receive data between the personal virtual environment synthesis device 15 and the terminal 11, the camera 12, and the microphone array 13.
  • Storage 153 includes, for example, a hard disk drive (HDD) or a solid state drive (SSD). Storage 153 stores information used in various processes in the personal virtual environment synthesis device 15.
  • HDD hard disk drive
  • SSD solid state drive
  • Drive 154 is a device for reading software stored in storage medium 155.
  • Drive 154 includes, for example, a CD (Compact Disk) drive or a DVD (Digital Versatile Disk) drive.
  • CD Compact Disk
  • DVD Digital Versatile Disk
  • the storage medium 155 is a medium that stores software electrically, magnetically, optically, mechanically, or chemically.
  • the storage medium 155 may store programs for executing various processes in the personal virtual environment synthesis device 15.
  • the personal virtual environment synthesis device 15 functions as a computer including a voice extraction unit 21, a conversational person estimation unit 22, a personal characteristic acquisition unit 23, a storage unit 24, a personal situation acquisition unit 25, a concentration level acquisition unit 26, an environment control unit 27, and an output unit 28.
  • the audio extraction unit 21 is configured to be able to separate and extract audio data of multiple conversations in space A and background sound data based on the audio data acquired by the microphone array 13, for example.
  • the interlocutor estimation unit 22 estimates the users U participating in each conversation. For example, the interlocutor estimation unit 22 estimates the user U (speaker) who is speaking and the position of the speaker using the image capture data of the camera 12 and the audio data of multiple conversations extracted by the audio extraction unit 21. The interlocutor estimation unit 22 also estimates the speaker's interlocutor and other users U (listeners of the conversation, etc.) who will participate in the conversation based on the face and line of sight of each user U using the image capture data of the camera 12, for example. Note that hereinafter, the speaker, the speaker's interlocutor, and other users U who will participate in the conversation are collectively referred to as interlocutors.
  • the personal characteristic acquisition unit 23 generates personal characteristic information 241 based on data acquired from the terminal 11, for example.
  • the personal characteristic information 241 includes, for example, the optimum value and boundary value for the window size ratio (percentage) of the terminal 11 in the business of each user U, and the volume ratio (percentage) of various audio data when synthesizing the audio data of the virtual environmental sound output from the headphones 14.
  • the window size ratio is, for example, the ratio of the size of the active window to the maximum value of the size (height and width, etc.) of the active window of the terminal 11.
  • the volume ratio is, for example, the ratio of the volume of each audio data to the maximum volume of the audio data.
  • the volume ratio includes, for example, the volume ratio of conversation sound data (conversation volume ratio), the volume ratio of background sound data (background volume ratio), and the volume ratio of notification sound data (notification volume ratio).
  • the optimum value of the window size ratio is, for example, a lower limit value that is estimated to need to be set to or above the value in order to make the business performance of each user U equal to or above the standard performance.
  • the boundary value of the window size ratio is, for example, an upper limit value at which the business performance of each user U is estimated to deteriorate significantly when the window size ratio is equal to or less than the boundary value.
  • the optimal value of the window size ratio can be higher than the boundary value of the window size ratio.
  • the optimal value of the volume ratio is, for example, an upper limit value that is estimated to be required to be set to or less than the value in order to make the business performance of each user U equal to or greater than the standard performance.
  • the boundary value of the volume ratio is, for example, a lower limit value at which the business performance of each user U is estimated to deteriorate significantly when the volume ratio is equal to or greater than the boundary value.
  • the optimal value of the volume ratio can be lower than the boundary value of the volume ratio.
  • each of the window size ratio and the volume ratio is also referred to as a control ratio.
  • increasing the window size ratio and decreasing the volume ratio are also referred to as increasing the occlusion rate of the surrounding environment.
  • decreasing the window size ratio and increasing the volume ratio are also referred to as decreasing the occlusion rate of the surrounding environment.
  • the personal characteristic acquisition unit 23 may generate the personal characteristic information 241 based on data acquired by, for example, a built-in camera and a biosensor of the terminal 11.
  • the biosensor is a sensor that can be worn on the body and is configured to acquire data based on, for example, pulse waves, body temperature, and skin conductance.
  • the storage unit 24 is a memory area of the storage 153.
  • the storage unit 24 stores, for example, personal characteristic information 241 and setting information 242.
  • FIG. 5 is a diagram showing an example of personal characteristic information according to an embodiment.
  • the personal characteristic information 241 for example, the optimum value and boundary value of the window size ratio, the optimum value and boundary value of the conversation volume ratio, the optimum value and boundary value of the notification volume ratio, and the optimum value and boundary value of the background volume ratio are stored for each user U.
  • setting information 242 default values and the latest setting values for the window size ratio, conversation volume ratio, notification volume ratio, and background volume ratio are stored for each user U.
  • the setting values are updated, for example, each time the window size ratio, conversation volume ratio, notification volume ratio, and background volume ratio are calculated.
  • the personal situation acquisition unit 25 generates data on the situation of each user U, for example, using the speaker and the position of the speaker estimated by the interlocutor estimation unit 22, and the image capture data of the camera 12. More specifically, the personal situation acquisition unit 25 detects at least one of the following: that each user U is not gazing at his/her own terminal 11 and has turned his/her face away from the terminal 11 (the action of each user turning their face away from the terminal 11), and that the terminal 11 used by the user U is in an unoperated state. Then, the personal situation acquisition unit 25 generates data on the time at which these are detected, for example. The personal situation acquisition unit 25 may acquire the image capture data of the camera 12 via the interlocutor estimation unit 22 or directly from the camera 12.
  • the personal situation acquisition unit 25 may generate data regarding the situation of each user U based on data acquired by the built-in camera and biosensor of the terminal 11.
  • the personal status acquisition unit 25 is also configured to be able to determine whether or not user U's work has been completed, for example, based on the generated data relating to the status of each user U and the captured data of the camera 12.
  • the personal status acquisition unit 25 may be configured to acquire information relating to the state of the terminal 11 (power on/off state, etc.). In this case, the personal status acquisition unit 25 may be configured to determine whether or not user U's work has been completed, for example, based on the acquired information relating to the state of the terminal 11.
  • the concentration level acquisition unit 26 calculates a concentration level indicating whether each user U is concentrating on work based on data related to the status of each user U generated by the personal status acquisition unit 25.
  • the concentration level acquisition unit 26 assumes that the concentration level is increasing as the interval between the time when at least one of the face-averting action and the no-operation state is detected increases.
  • the concentration level acquisition unit 26 assumes that the concentration level is decreasing as the interval between the time when at least one of the face-averting action and the no-operation state is detected decreases. More specifically, the concentration level acquisition unit 26 may assume that the concentration level is increasing as the interval between the time when the face is averted from the terminal 11 increases, or as the time when each user U operates the terminal 11 is longer and the time when no operation is shorter.
  • the concentration level acquisition unit 26 may assume that the concentration level is decreasing as the interval between the time when the face is averted increases, or as the time when the operation is shorter and the time when no operation is longer.
  • the concentration level acquisition unit 26 may calculate the concentration level from the blinking and facial expressions of each user U using data acquired by a camera provided in the terminal 11.
  • the concentration level acquisition unit 26 may also calculate the concentration level using information from the camera 12, the built-in camera of the terminal 11, and a biosensor. In this case, the concentration level acquisition unit 26 may calculate the concentration level by combining and weighting values calculated using data acquired by the camera 12, the camera provided in the terminal 11, and the biosensor based on the characteristics of each user U measured in advance.
  • the environment control unit 27 determines, for example, a control characteristic (control function) for calculating a window size ratio, a conversation volume ratio, a notification volume ratio, and a background volume ratio according to the concentration level of each user U based on the personal characteristic information 241.
  • the environment control unit 27 controls the occlusion rate of the surrounding environment for each user U by calculating a control ratio according to the calculated concentration level of each user U based on the control function.
  • the environment control unit 27 also updates, for example, the setting value of the currently applied control ratio stored in the setting information 242 to the calculated control ratio.
  • the environment control unit 27 is configured, for example, to increase the occlusion rate of the conversation sound data, background sound data, and notification sound data (lower the control ratio and lower the volume) as the concentration level of each user U increases.
  • the environment control unit 27 is also configured, for example, to decrease the occlusion rate of the conversation sound data, background sound data, and notification sound data (higher the control ratio and raise the volume) as the concentration level of each user U decreases.
  • the environment control unit 27 is configured to, for example, increase the occlusion rate of the window size (increase the control ratio and make the active window larger) as the concentration level of each user U increases.
  • the environment control unit 27 is configured to, for example, decrease the occlusion rate of the window size (decrease the control ratio and make the active window smaller) as the concentration level of each user U decreases.
  • the surrounding environment of each user U controlled by the environment control unit 27 may be a combination of at least one of conversation sound data, background sound data, notification sound data, and the active window size of the terminal 11, based on the characteristics of the user U measured in advance.
  • the output unit 28 synthesizes the audio data of the virtual environmental sound based on the conversation sound data and background sound data extracted by the audio extraction unit 21, the notification sound data, the information of the interlocutor estimated by the interlocutor estimation unit 22, and the control ratio of various audio data stored in the setting information 242.
  • the output unit 28 may be configured to use notification sound data stored in advance when there is a notification to the terminal 11.
  • the output unit 28 also outputs the audio data of the synthesized virtual environmental sound to the headphones 14.
  • the output unit 28 may, for example, make the audio data of the virtual environmental sound for each user U audio data in which a conversation in which the user U participates among multiple conversations is emphasized and a conversation in which the user U does not participate is removed or attenuated.
  • the output unit 28 also controls the terminal 11 to change the window size based on the window size ratio generated by the environment control unit 27.
  • the personal virtual environment synthesis device 15 controls the window size of the terminal 11 and the audio data of the virtual environment sound output to the headphones 14 according to the concentration level of each user U.
  • FIG. 6 is a flowchart showing an example of a personal virtual environment synthesis processing in the personal virtual environment synthesis device according to the embodiment.
  • each user U executes a test task (start).
  • each user U performs the test task while changing the window size of the terminal 11 and the volumes of the conversation sound data, background sound data, and notification sound data.
  • the personal situation acquisition unit 25 generates characteristics of each user U for the window size of the terminal 11, the conversation sound data, background sound data, and notification sound data output from the headphones 14 (S11). That is, the personal characteristic acquisition unit 23 estimates optimal values and boundary values for each of the window size, conversation sound volume ratio, background sound volume ratio, and notification sound volume ratio.
  • the test task includes, for example, a test to determine whether the user U notices the notification sound output from the headphones 14.
  • the personal characteristic acquisition unit 23 may generate personal characteristic information 241 in advance prior to the personal virtual environment synthesis process.
  • the voice extraction unit 21 starts to separate and extract the voice data of the multiple conversations and the background sound data.
  • the conversational party estimation unit 22 estimates conversational party information using the image data of the camera 12 and the voice data of the multiple conversations extracted by the voice extraction unit 21. Note that if the personal characteristic information 241 has been generated in advance, the operation of the personal virtual environment synthesis device 15 starts when user U starts work.
  • the output unit 28 also starts synthesizing the audio data of the virtual environmental sound using, for example, a default control ratio, the audio data of multiple conversations and background sound data extracted by the audio extraction unit 21, the notification sound data, and the information of the conversational party estimated by the conversational party estimation unit 22. That is, the output unit 28 synthesizes the audio data of the virtual environmental sound using, for example, the conversation sound data, background sound data, and notification sound data having a volume corresponding to the default control ratio. Then, the output unit 28 starts outputting the synthesized audio data to the headphones 14.
  • the size of the active window is also determined based on the default control ratio.
  • the personal situation acquisition unit 25 generates data on the situation of each user U when acquiring the personal situation (S12).
  • the personal situation acquisition unit 25 detects an interruption in the work of each user U (a break in the concentration of the user U) based on, for example, the speaker estimated by the interlocutor estimation unit 22, the position of the speaker, and the photographic data of the camera 12. More specifically, the personal situation acquisition unit 25 detects, as an interruption in work, at least one of the actions of each user U turning their face away from the terminal 11 and the terminal 11 used by the user U being in an unoperated state.
  • the personal situation acquisition unit 25 also generates data on, for example, the time when the user U interrupted their work.
  • the concentration level acquisition unit 26 calculates the concentration level of each user U based on the individual status acquired in S12 (data on the time when the user U stopped working) (S13). The process of calculating the concentration level will be described later.
  • the environment control unit 27 calculates the control ratios for the window size of the terminal 11 and the volume of various types of audio data used to synthesize audio data of the virtual environmental sound, according to the concentration level of each user U calculated in S13.
  • the environment control unit 27 also updates the various control ratios of each user stored in the setting information 242. The process of calculating the control ratios will be described later. In this manner, the control of the occlusion rate is executed (S14).
  • the output unit 28 starts synthesizing the audio data of the virtual environmental sound by using the updated control ratios for the volume of the various audio data, the audio data of the multiple conversations and background sound data extracted by the audio extraction unit 21, the notification sound data, and the information of the conversational parties estimated by the conversational party estimation unit 22 (S15).
  • the output unit 28 synthesizes the audio data of the virtual environmental sound by using, for example, the conversation sound data, background sound data, and notification sound data having a volume corresponding to the updated control ratio. Then, the output unit 28 outputs the synthesized audio data to the headphones 14. The output unit 28 also changes the window size of each terminal 11 based on the updated window size ratio. In this manner, the output unit 28 synthesizes the personal environment.
  • the personal status acquisition unit 25 determines whether all users U have completed their tasks (S16).
  • the above-described personal virtual environment synthesis process is applied, for example, to tasks that are performed in real time.
  • Fig. 7 is a diagram for explaining the time at which a user interrupts a task, as detected in the personal virtual environment synthesis device according to the embodiment.
  • the time at which it is detected that the user U interrupted the task for the nth time since starting the task is indicated as time T(fout(n)).
  • the concentration level acquisition unit 26 calculates the rate of change in the concentration level, expressed by the following formula (1), based on the data on the time when each user U, shown in FIG. 7, interrupted their work.
  • the concentration acquisition unit 26 converts the rate of change in concentration fd(n) into a value fs(n) between 0 and 1 using the sigmoid function expressed by the following formula (2).
  • the concentration acquisition unit 26 calculates the concentration fr(n) at time T(fout(n)) based on the above-mentioned value fs(n) using the geometric mean expressed by the following formula (3).
  • Fig. 8 is a graph for explaining an example of the relationship between the concentration level and the control ratio in the personal virtual environment synthesis processing according to the embodiment.
  • the concentration level is shown on the horizontal axis
  • the control ratio is shown on the vertical axis.
  • the environment control unit 27 determines, for each user U, a control function between the concentration level and each of the window size ratio, conversation volume ratio, notification volume ratio, and background volume ratio, based on, for example, the personal characteristic information 241.
  • control functions for the window size ratio, conversation volume ratio, notification volume ratio, and background volume ratio are shown.
  • the present invention is not limited to these, and each of the window size ratio, conversation volume ratio, notification volume ratio, and background volume ratio may have a different control function for each user U.
  • the environment control unit 27 calculates the window size ratio srw(n) (percentage) based on the concentration level fr(n) at time T(fout(n)) using, for example, the following formula (4). As shown in FIG. 8, the higher the concentration level fr(n), the larger the window size ratio srw(n).
  • the environment control unit 27 can determine the control function for the window size ratio based, for example, on the optimal value and boundary value of the window size ratio for each user U.
  • the environmental control unit 27 also calculates the conversation volume ratio src(n) (percentage) based on the concentration level fr(n) at time T(fout(n)) using, for example, the following formula (5). The higher the concentration level fr(n), the smaller the conversation volume ratio src(n).
  • the environmental control unit 27 can determine a control function for the conversation volume ratio based on, for example, the optimal value and boundary value of the conversation volume ratio for each user U.
  • the environmental control unit 27 also calculates the notification volume ratio srn(n) (percentage) based on the concentration level fr(n) at time T(fout(n)) using, for example, the following formula (6). As shown in FIG. 8, the higher the concentration level fr(n), the smaller the notification volume ratio srn(n).
  • the environmental control unit 27 can determine a control function for the notification volume ratio based on, for example, the optimal value and boundary value of the notification volume ratio for the conversation sound data of each user U.
  • the environment control unit 27 can calculate the background volume ratio srb(n) (percentage) using, for example, an equation similar to the above equation (6). For the background volume ratio as well, the environment control unit 27 can determine a control function for the background volume ratio based on, for example, the optimal value and boundary value of the background volume ratio for each user U.
  • the personal virtual environment synthesis device 15 includes a concentration level acquisition unit 26 that calculates the concentration level of the user U, and an environment control unit 27 that changes the shading rate of the user U with respect to the surrounding environment based on the calculated concentration level.
  • the personal virtual environment synthesis device 15 is configured to increase the shading rate of the surrounding environment of a user U with a high concentration level, for example.
  • the personal virtual environment synthesis device 15 is configured to decrease the shading rate of the surrounding environment of a user U with a low concentration level, for example.
  • the personal virtual environment synthesis device 15 can gradually change the shading rate of the surrounding environment according to the concentration level of each user U. Therefore, the user U can naturally increase his/her concentration, or naturally relax his/her tension when his/her concentration level decreases. Therefore, it is possible to suppress excessive accumulation of fatigue in the user U, and to encourage the user U to perform work without strain.
  • control function for calculating the control ratio (controlling the shading rate) for each user U is determined, for example, based on the characteristics (personal characteristic information 241) of the user U that are acquired in advance.
  • the personal virtual environment synthesis device 15 can control the surrounding environment in a manner suitable for each user U.
  • the program that executes the personal virtual environment synthesis process is executed by the personal virtual environment synthesis device 15, but this is not limited to the above.
  • the program that executes the personal virtual environment synthesis process may be executed by computing resources on the cloud.
  • the present invention is not limited to the above-described embodiments, and can be modified in various ways during implementation without departing from the gist of the invention.
  • the embodiments may also be implemented in appropriate combination, in which case the combined effects can be obtained.
  • the above-described embodiments include various inventions, and various inventions can be extracted by combinations selected from the multiple constituent elements disclosed. For example, if the problem can be solved and an effect can be obtained even if some constituent elements are deleted from all the constituent elements shown in the embodiments, the configuration from which these constituent elements are deleted can be extracted as an invention.

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • User Interface Of Digital Computer (AREA)

Abstract

一実施形態の個人仮想環境合成装置は、ユーザの集中度を算出する集中度取得部と、算出された上記集中度に基づいて、上記ユーザの周辺環境に対する遮蔽率の制御を行う制御部と、を備える。

Description

個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム
 実施形態は、個人仮想環境合成装置、個人仮想環境合成方法、及びプログラムに関する。
 視聴者が会話の構造等を理解しやすくし、自動的に仮想空間映像の視点が切り替わっていくようにする技術が研究されている。例えば、特許文献2に記載の技術では、観測装置から得られる映像を用いて、会話参加者の顔の位置、姿勢、及び視線が推定される。また、観測装置から得られる音声を用いて、発話の有無及び到来方向が推定される。以上のようにして推定された顔の位置、姿勢、視線、発話の有無、及び到来方向を用いて、仮想カメラを注視する度合いである注視度、仮想空間の原点に対する仮想カメラの方位角、及び仮想カメラの視点を制御する制御パラメータが求められる。そして、会話参加者の画像が部分平面に射影され、射影された部分平面が実際の会話参加者の配置と対応するように仮想空間上の水平面に配置される。それから、制御パラメータを用いて、仮想カメラの視点の仮想空間映像が生成される。
 また、音と知的作業の効率との関係について、非特許文献1に記載されるように、作業内容に関連した有意騒音が作業に影響を与えることが考えられている。また、同室内で交わされる会話が自分にも関係することを考慮することが集中を妨げる要因となる可能性が、考えられている。
 そこで、同室内で複数の会話が発生する場合等に、個人ごとに、必要な音及び不要な音を仮想的な環境音として合成する技術が研究されている。このような技術において、複数の会話の各々を分離して密閉することで、各会話に参加する人が、会話に集中できるようにされる。また、会話に参加しない個人に対しては、会話を聞こえにくくすることで、作業に集中できる音環境が提供される。
 例えば、特許文献1に記載の技術では、マイクロフォンアレイ及び全方位カメラのデータから、話者の音声と話者の位置とが推定される。また、例えば、全方位カメラのデータにおける話者の顔や視線の向きから、話者が対面する個人が推定される。また、例えば、全方位カメラのデータにおける話者の顔や視線の向きから、各個人が対面する個人が推定される。以上のような推定に基づき、室内の会話状況が管理される。そして、各個人に必要な会話を強調し、不要な会話を除去または減衰させた仮想環境音が合成される。それから、合成された仮想環境音が、各個人へ送信される。
特開2021-189247号公報 特開2010-191544号公報 特開2018-159979号公報
澤木 美奈子,山森 和彦,"騒音・BGMが知的作業に与える影響",騒音制御(journal of INCE/J),第16巻,第5号,239-242ページ,1992年10月
 しかしながら、会話に参加しない個人に対して会話を聞こえにくくする場合、各個人の集中力及び業務の効率は向上するが、過集中となることで、疲労蓄積による体調不良、及び業務への支障などが引き起こされる可能性がある。
 過集中となってしまうことを抑制するため、特許文献3に記載されるような、業務のスケジュールや業務の経過時間に応じて休憩を勧めるよう通知する技術を適用することが考えられる。しかしながら、この場合、各個人の集中力が高まったタイミングに休憩を勧める通知がされることで、業務の妨げとなる可能性がある。
 本発明は、上記事情に着目してなされたもので、その目的とするところは、業務を妨げることなく、自然に集中力を高めたり、緊張を緩めたりすることで、各個人の過度な疲労蓄積を抑制する手段を提供することにある。
 一態様の個人仮想環境合成装置は、一実施形態の個人仮想環境合成装置は、ユーザの集中度を算出する集中度取得部と、算出された上記集中度に基づいて、上記ユーザの周辺環境に対する遮蔽率の制御を行う制御部と、を備える。
 実施形態によれば、自然に集中力を高めたり、緊張を緩めたりすることで、各個人の過度な疲労蓄積を抑制する手段を提供することができる。
図1は、実施形態に係る個人仮想環境合成システムの一例を示す図である。 図2は、実施形態に係る端末のハードウェア構成の一例を示すブロック図である。 図3は、実施形態に係る個人仮想環境合成装置のハードウェア構成の一例を示すブロック図である。 図4は、実施形態に係る個人仮想環境合成装置の機能構成の一例を示すブロック図である。 図5は、実施形態に係る個人特性情報の一例を示す図である。 図6は、実施形態に係る個人仮想環境合成装置における個人仮想環境合成処理の一例を示すフローチャートである。 図7は、実施形態に係る個人仮想環境合成装置において検出される、業務が中断された時刻を説明するための図である。 図8は、実施形態に係る個人仮想環境合成における集中度と制御比率との関係の一例を説明するためのグラフである。
 以下、図面を参照して実施形態について説明する。なお、以下の説明において、同一の機能及び構成を有する構成要素については、共通する参照符号を付す。
 1. 実施形態
 1.1 個人仮想環境合成システム
 図1は、実施形態に係る個人仮想環境合成システムの一例を示す図である。
 個人仮想環境合成システム1は、例えば、同じ空間A内で業務を行う複数のユーザの各々の仮想環境を合成するシステムである。個人仮想環境合成システム1は、端末11-1、11-2、11-3、11-4、11-5、及び11-6、カメラ12、マイクロフォンアレイ13、ヘッドフォン14-1、14-2、14-3、14-4、14-5、及び14-6、並びに個人仮想環境合成装置15を含む。なお、以下の説明において、端末11-1、11-2、11-3、11-4、11-5、及び11-6の各々を区別しない場合には、端末11-1、11-2、11-3、11-4、11-5、及び11-6の各々を単に端末11と呼ぶ。また、ヘッドフォン14-1、14-2、14-3、14-4、14-5、及び14-6の各々を区別しない場合には、ヘッドフォン14-1、14-2、14-3、14-4、14-5、及び14-6の各々を単にヘッドフォン14と呼ぶ。
 各端末11は、パーソナルコンピュータ又はスマートフォン等である。端末11と個人仮想環境合成装置15とは、例えば、ネットワークNWを介して互いに通信可能に構成される。端末11-1、11-2、11-3、11-4、11-5、及び11-6はそれぞれ、ユーザU1、U2、U3、U4、U5、及びU6の業務に使用される。ユーザU1、U2、U3、U4、U5、及びU6は、同じ空間A内において業務を行う。なお、以下の説明において、ユーザU1、U2、U3、U4、U5、及びU6の各々を区別しない場合には、ユーザU1、U2、U3、U4、U5、及びU6の各々を単にユーザUと呼ぶ。
 カメラ12は、例えば、ユーザUの撮影データを取得する。カメラ12は、例えば、全方位カメラである。カメラ12は、例えば、ネットワークNWを介して、取得した撮影データを個人仮想環境合成装置15に送信可能に構成される。
 マイクロフォンアレイ13は、複数のマイクロフォンMPを含む。マイクロフォンアレイ13は、複数のマイクロフォンMPを用いて空間A内の音声データを取得する。空間A内の音声データは、例えば、ユーザU1~U6による会話の音声データ(会話音データ)、及び背景音の音声データ(背景音データ)を含む。会話音データは、ユーザU1~U6による複数の会話の音声データを含み得る。背景音データは、例えば、空調やプリンタの作動音等のデータを含む。マイクロフォンアレイ13は、例えば、ネットワークNWを介して、取得した音声データを個人仮想環境合成装置15に送信する。
 ヘッドフォン14-1、14-2、14-3、14-4、14-5、及び14-6はそれぞれ、ユーザU1、U2、U3、U4、U5、及びU6によって使用される。各ヘッドフォン14には、当該ヘッドフォン14を使用するユーザU用の音声データが、例えば、当該ユーザUが使用する端末11を介して、個人仮想環境合成装置15から送信される。なお、各ヘッドフォン14には、音声データが、個人仮想環境合成装置15から直接的に送信されてもよい。
 個人仮想環境合成装置15は、例えば、カメラ12の撮影データ、マイクロフォンアレイ13により取得される音声データ、及びヘッドフォン14に出力する音声データを管理するサーバである。個人仮想環境合成装置15は、カメラ12の撮影データ、及びマイクロフォンアレイ13により取得される音声データに基づき、例えば、会話音データ、及び背景音データをそれぞれ分離して抽出可能に構成される。また、個人仮想環境合成装置15は、例えば、カメラ12から取得した映像に基づき、各ユーザUの集中度(パーセント)を算出する。そして、個人仮想環境合成装置15は、各ユーザUの集中度に基づき、上述のように分離して抽出された各音声データ、及び端末11への通知の音声データ(通知音データ)を用いて、当該ユーザU用の音声データ(仮想環境音の音声データ)を合成する。通知音データは、例えば、端末11に通知されるメールの着信音等のデータを含む。それから、個人仮想環境合成装置15は、仮想環境音の音声データを、上記ユーザUのヘッドフォン14に出力する。また、個人仮想環境合成装置15は、例えば、各ユーザUの集中度に基づき、当該ユーザUの使用する端末11上の作業ウィンドウ(アクティブウィンドウ)のサイズを制御する。なお、個人仮想環境合成装置15は、空間Aの内側にあっても、空間Aの外側にあってもよく、個人仮想環境合成装置15がある場所は特に限定されない。
 1.2 端末
 次に、実施形態に係る端末の構成について説明する。
 図2は、実施形態に係る端末のハードウェア構成の一例を示すブロック図である。図2に示すように、各端末11は、制御回路111、通信モジュール112、及びユーザインタフェース113を含む。端末11-1、11-2、11-3、11-4、11-5、及び11-6は、互いに同等の構成を有する。
 制御回路111は、端末11の各構成要素を全体的に制御する回路である。制御回路111は、CPU(Central Processing Unit)、RAM(Random Access Memory)、及びROM(Read Only Memory)等を含む。制御回路111のROMは、端末11における各種処理で使用されるプログラム等を記憶する。制御回路111のCPUは、制御回路111のROMに記憶されるプログラムにしたがって、端末11の全体を制御する。制御回路111のRAMは、制御回路111のCPUの作業領域として使用される。
 通信モジュール112は、端末11と個人仮想環境合成装置15との間のデータの送受信に使用される回路である。
 ユーザインタフェース113は、ユーザUと制御回路111との間の通信を司るインタフェースである。ユーザインタフェース113は、入力機器及び出力機器を含む。入力機器は、例えば、音声マイク、タッチパネル、及び操作ボタン等を含む。出力機器は、例えば、スピーカ、及びディスプレイ等を含む。
 なお、図示しないが、各端末11はカメラ(内蔵カメラ)を含んでもよい。
 1.3 個人仮想環境合成装置
 次に、実施形態に係る個人仮想環境合成装置15の構成について説明する。
 1.3.1 ハードウェア構成
 図3は、実施形態に係る個人仮想環境合成装置のハードウェア構成の一例を示すブロック図である。図3に示すように、個人仮想環境合成装置15は、制御回路151、通信モジュール152、ストレージ153、ドライブ154、及び記憶媒体155を含む。
 制御回路151は、個人仮想環境合成装置15の各構成要素を全体的に制御する回路である。制御回路151は、CPU、RAM、及びROM等を含む。制御回路151のROMは、個人仮想環境合成装置15における各種処理で使用されるプログラム等を記憶する。制御回路151のCPUは、制御回路151のROMに記憶されるプログラムにしたがって、個人仮想環境合成装置15の全体を制御する。制御回路151のRAMは、制御回路151のCPUの作業領域として使用される。
 通信モジュール152は、個人仮想環境合成装置15と、端末11、カメラ12、及びマイクロフォンアレイ13との間のデータの送受信に使用される回路である。
 ストレージ153は、例えば、HDD(Hard Disk Drive)又はSSD(Solid State Drive)を含む。ストレージ153は、個人仮想環境合成装置15における各種処理で使用される情報が記憶される。
 ドライブ154は、記憶媒体155に記憶されたソフトウェアを読み込むための機器である。ドライブ154は、例えば、CD(Compact Disk)ドライブ又はDVD(Digital Versatile Disk)ドライブを含む。
 記憶媒体155は、ソフトウェアを、電気的、磁気的、光学的、機械的又は化学的作用によって記憶する媒体である。記憶媒体155は、個人仮想環境合成装置15における各種処理を実行するためのプログラムを記憶してもよい。
 1.3.2 機能構成
 図4は、実施形態に係る個人仮想環境合成装置の機能構成の一例を示すブロック図である。制御回路151のCPUは、制御回路151のROM又は記憶媒体155に記憶されたプログラムを制御回路151のRAMに展開する。そして、制御回路151のCPUは、制御回路151のRAMに展開されたプログラムを解釈及び実行する。これにより、個人仮想環境合成装置15は、音声抽出部21、会話者推定部22、個人特性取得部23、記憶部24、個人状況取得部25、集中度取得部26、環境制御部27、及び出力部28を備えるコンピュータとして機能する。
 音声抽出部21は、例えば、マイクロフォンアレイ13によって取得された音声データに基づき、空間A内における複数の会話の音声データ、及び背景音データをそれぞれ分離して抽出可能に構成される。
 会話者推定部22は、例えば、各会話に参加するユーザUを推定する。会話者推定部22は、例えば、カメラ12の撮影データ、及び音声抽出部21によって抽出された複数の会話の音声データを用いて、発話するユーザU(話者)、及び話者の位置を推定する。また、会話者推定部22は、例えば、カメラ12の撮影データを用いて、各ユーザUの顔や視線に基づき、話者の話し相手、及び会話に参加するその他のユーザU(会話の聞き手等)を推定する。なお、以下では、話者、話者の話し相手、及び会話に参加するその他のユーザUを総称して会話者と呼ぶ。
 個人特性取得部23は、例えば、端末11から取得するデータに基づき、個人特性情報241を生成する。個人特性情報241は、例えば、各ユーザUの業務における端末11のウィンドウサイズ比率(パーセント)、及びヘッドフォン14から出力される仮想環境音の音声データを合成する際の各種音声データの音量比率(パーセント)についての最適値及び境界値を含む。ウィンドウサイズ比率は、例えば、端末11のアクティブウィンドウのサイズ(高さ及び幅等)の最大値に対する、アクティブウィンドウのサイズの比率である。音量比率は、例えば、各音声データの音量の最大値に対する、当該音声データの音量の比率である。音量比率は、例えば、会話音データの音量比率(会話音量比率)、背景音データの音量比率(背景音量比率)、及び通知音データの音量比率(通知音量比率)を含む。ウィンドウサイズ比率の最適値は、例えば、各ユーザUの業務の成績を基準の成績以上にするために当該値以上にする必要があると推定される下限値である。また、ウィンドウサイズ比率の境界値は、例えば、ウィンドウサイズ比率が当該境界値以下である場合に、各ユーザUの業務の成績が顕著に悪化すると推定される上限値である。ウィンドウサイズ比率の最適値は、ウィンドウサイズ比率の境界値より高くなり得る。音量比率の最適値は、例えば、各ユーザUの業務の成績を基準の成績以上にするために当該値以下にする必要があると推定される上限値である。また、音量比率の境界値は、例えば、音量比率が当該境界値以上である場合に、各ユーザUの業務の成績が顕著に悪化すると推定される下限値である。音量比率の最適値は、音量比率の境界値より低くなり得る。なお、以下の説明において、ウィンドウサイズ比率及び音量比率を区別しない場合には、ウィンドウサイズ比率及び音量比率の各々を制御比率とも呼ぶ。また、以下の説明では、ウィンドウサイズ比率を高くすること、及び音量比率を低くすることを、周辺環境の遮蔽率を高くするともいう。また、ウィンドウサイズ比率を低くすること、及び音量比率を高くすることを、周辺環境の遮蔽率を低くするともいう。
 なお、図示しないが、個人特性取得部23は、例えば、端末11の内蔵カメラ、及び生体センサによって取得されるデータに基づいて、個人特性情報241を生成してもよい。なお、生体センサは、例えば、脈波、体温、皮膚伝導に基づくデータを取得するように構成された、身体に装着可能なセンサである。
 記憶部24は、ストレージ153のメモリ領域である。記憶部24は、例えば、個人特性情報241、及び設定情報242を記憶する。
 図5は、実施形態に係る個人特性情報の一例を示す図である。図5に示されるように、個人特性情報241において、例えば、ウィンドウサイズ比率の最適値及び境界値、会話音量比率の最適値及び境界値、通知音量比率の最適値及び境界値、並びに背景音量比率の最適値及び境界値が、ユーザUごとに記憶される。
 設定情報242には、ウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率についてデフォルト値、及び最新の設定値が、ユーザUごとに記憶される。当該設定値は、例えば、ウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率が算出されるごとに更新される。
 図4に戻り、個人仮想環境合成装置15の機能構成について説明を続ける。
 個人状況取得部25は、例えば、会話者推定部22によって推定された話者、及び当該話者の位置、並びにカメラ12の撮影データを用いて、各ユーザUの状況に関するデータを生成する。より具体的には、個人状況取得部25は、例えば、各ユーザUが自身の端末11を注視せず、端末11から顔をそらすこと(各ユーザが端末11から顔をそらす動作)、及び当該ユーザUが使用する端末11が無操作状態であることのうち少なくともいずれか1つを検出する。そして、個人状況取得部25は、例えば、これらを検出した時刻のデータを生成する。なお、個人状況取得部25は、会話者推定部22を介して、又はカメラ12から直接的に、カメラ12の撮影データを取得し得る。
 なお、以下では、個人状況取得部25がカメラ12の撮影データを用いて各ユーザUの状況に関するデータを生成する例が説明されるが、これに限られない。個人状況取得部25は、例えば、個人特性取得部23と同様に、端末11の内蔵カメラ、及び生体センサによって取得されるデータに基づいて、各ユーザUの状況に関するデータを生成してもよい。
 また、個人状況取得部25は、例えば、生成した各ユーザUの状況に関するデータ、及びカメラ12の撮影データ等に基づき、ユーザUの業務が終了したかどうかを判定可能に構成される。個人状況取得部25は、例えば、端末11の状態(電源のオン状態及びオフ状態等)に関する情報を取得するように構成されていてもよい。この場合、個人状況取得部25は、例えば、取得した端末11の状態に関する情報に基づき、ユーザUの業務が終了したかどうかを判定するように構成されてもよい。
 集中度取得部26は、個人状況取得部25によって生成された各ユーザUの状況に関するデータに基づき、各ユーザUが業務に集中しているかどうかを示す集中度を算出する。集中度取得部26は、例えば、顔をそらす動作、及び無操作状態のうち少なくともいずれか1つを検出する時刻の間隔が広がるほど、集中度が上昇しているとする。また、集中度取得部26は、例えば、顔をそらす動作、及び無操作状態のうち少なくともいずれか1つを検出する時刻の間隔が狭まるほど、集中度が低下しているとする。より具体的に、集中度取得部26は、例えば、端末11から顔をそらす時刻の間隔が広がるほど、又は各ユーザUの端末11の有操作時間が長く、無操作時間が短くなるほど、集中度が上昇しているとし得る。また、集中度取得部26は、例えば、顔をそらす時刻の間隔が狭まるほど、又は有操作時間が短く、無操作時間が長くなるほど、集中度が低下しているとし得る。なお、集中度取得部26は、例えば、端末11が備えるカメラによって取得されるデータを用いて、各ユーザUの瞬きや表情から、集中度を算出してもよい。また、集中度取得部26は、カメラ12、端末11の内蔵カメラ、生体センサからの情報を用いて、集中度を算出してもよい。この場合、集中度取得部26は、予め測定されたユーザUごとの特性に基づき、カメラ12、端末11が備えるカメラ、生体センサにより取得されるデータを用いて算出される値の組み合わせ、及び重みづけをすることで集中度を算出し得る。
 環境制御部27は、例えば、個人特性情報241に基づいて、各ユーザUの集中度に応じたウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率を算出するための制御特性(制御関数)を決定する。環境制御部27は、制御関数に基づいて、算出された各ユーザUの集中度に応じた制御比率を算出することで、各ユーザUについての周辺環境の遮蔽率を制御する。また、環境制御部27は、例えば、設定情報242に記憶される現在適用される制御比率の設定値を、算出した制御比率に更新する。環境制御部27は、例えば、各ユーザUの集中度が高いほど、会話音データ、背景音データ、及び通知音データの遮蔽率を高くする(制御比率を低くして、音量を下げる)ように構成される。また、環境制御部27は、例えば、各ユーザUの集中度が下がるほど、会話音データ、背景音データ、及び通知音データの遮蔽率を低くする(制御比率を高くして、音量を上げる)ように構成される。また、環境制御部27は、例えば、各ユーザUの集中度が高まるほど、ウィンドウサイズの遮蔽率を高くする(制御比率を高くして、アクティブウィンドウを大きくする)ように構成される。また、環境制御部27は、例えば、各ユーザUの集中度が下がるほど、ウィンドウサイズの遮蔽率を低くする(制御比率を低くして、アクティブウィンドウを小さくする)ように構成される。なお、環境制御部27により制御される各ユーザUの周辺環境は、事前に測定される当該ユーザUの特性に基づき、会話音データ、背景音データ、通知音データ、及び端末11のアクティブウィンドウサイズのうち少なくとも1つの組み合わせであってもよい。
 出力部28は、音声抽出部21によって抽出された会話音データ及び背景音データ、通知音データ、会話者推定部22によって推定された会話者の情報、並びに設定情報242に記憶される各種音声データの制御比率に基づいて、仮想環境音の音声データを合成する。出力部28は、仮想環境音の音声データの合成において、例えば、端末11への通知がある際に、予め記憶する通知音データを用いるように構成され得る。また、出力部28は、合成した仮想環境音の音声データを、ヘッドフォン14に出力する。なお、出力部28は、例えば、各ユーザU用の仮想環境音の音声データを、複数の会話のうち当該ユーザUが参加する会話が強調され、当該ユーザUが参加しない会話が除去又は減衰された音声データとし得る。また、出力部28は、環境制御部27によって生成されたウィンドウサイズ比率に基づき、ウィンドウサイズを変更するよう端末11を制御する。
 以上のような構成により、個人仮想環境合成装置15は、各ユーザUの集中度に応じて、端末11のウィンドウサイズ、及びヘッドフォン14に出力される仮想環境音の音声データを制御する。
 1.2 動作
 次に、実施形態に係る個人仮想環境合成装置の動作について説明する。
 1.2.1 個人仮想環境合成処理
 図6は、実施形態に係る個人仮想環境合成装置における個人仮想環境合成処理の一例を示すフローチャートである。
 まず、各ユーザUは、テストタスクを実行する(開始)。テストタスクの実行において、各ユーザUは、端末11のウィンドウサイズと、会話音データ、背景音データ、及び通知音データの音量と、を変更しながら、テストタスクを実施する。そして、個人状況取得部25は、端末11からのテストタスクの結果に基づいて、端末11のウィンドウサイズ、ヘッドフォン14から出力される会話音データ、背景音データ、及び通知音データについての、各ユーザUの特性を生成する(S11)。すなわち、個人特性取得部23は、ウィンドウサイズ、会話音量比率、背景音量比率、及び通知音量比率それぞれについての、最適値及び境界値を推定する。なお、テストタスクは、例えば、ヘッドフォン14から出力される通知音に、ユーザUが気付くかどうかを判定するテストを含む。また、個人特性取得部23は、個人仮想環境合成処理に先立ち、予め個人特性情報241を生成し得る。
 そして、個人特性情報241が生成された後、ユーザUの業務が開始すると、音声抽出部21が、複数の会話の音声データ及び背景音データをそれぞれ分離して抽出し始める。また、会話者推定部22が、カメラ12の撮影データ、及び音声抽出部21によって抽出された複数の会話の音声データを用いて、会話者の情報を推定する。なお、個人特性情報241が予め生成されている場合、個人仮想環境合成装置15の動作は、ユーザUの業務の開始により開始する。
 また、出力部28は、例えば、デフォルトの制御比率、音声抽出部21によって抽出された複数の会話の音声データ及び背景音データ、通知音データ、並びに会話者推定部22によって推定された会話者の情報を用いて仮想環境音の音声データの合成を開始する。すなわち、出力部28は、例えば、デフォルトの制御比率分の音量を有する会話音データ、背景音データ、及び通知音データを用いて、仮想環境音の音声データを合成する。そして、出力部28は、合成した音声データのヘッドフォン14への出力を開始する。また、アクティブウィンドウのサイズは、デフォルトの制御比率に基づいて決定される。
 個人状況取得部25は、個人状況の取得において、各ユーザUの状況に関するデータを生成する(S12)。個人状況取得部25は、例えば、会話者推定部22によって推定された話者、当該話者の位置、及びカメラ12の撮影データに基づき、各ユーザUの業務の中断(ユーザUの集中が途切れたこと)を検出する。より具体的に、個人状況取得部25は、例えば、各ユーザUが端末11から顔を背けた動作、及び当該ユーザUの使用する端末11が無操作状態であることのうち少なくともいずれか1つを、業務の中断であるとして検出する。また、個人状況取得部25は、例えば、ユーザUが業務を中断した時刻のデータを生成する。
 集中度取得部26は、S12において取得された個人状況(ユーザUが業務を中断した時刻のデータ)に基づき、各ユーザUの集中度を算出する(S13)。集中度を算出する処理については後述する。
 環境制御部27は、S13において算出された各ユーザUの集中度に応じて、端末11のウィンドウサイズ、及び仮想環境音の音声データの合成に用いられる各種音声データの音量についての制御比率を算出する。また、環境制御部27は、設定情報242に記憶される各ユーザの各種制御比率を更新する。制御比率を算出する処理については後述する。以上のようにして遮蔽率の制御が実行される(S14)
 出力部28は、更新された各種音声データの音量についての制御比率、音声抽出部21によって抽出された複数の会話の音声データ及び背景音データ、通知音データ、並びに会話者推定部22によって推定された会話者の情報を用いて仮想環境音の音声データの合成を開始する(S15)。すなわち、出力部28は、例えば、更新された制御比率分の音量を有する会話音データ、背景音データ、及び通知音データを用いて、仮想環境音の音声データを合成する。そして、出力部28は、合成した音声データをヘッドフォン14に出力する。また、出力部28は、更新されたウィンドウサイズ比率に基づき、各端末11のウィンドウサイズを変更する。以上のようにして、出力部28は、個人環境を合成する。
 個人状況取得部25は、全てのユーザUの業務が終了したかどうかを判定する(S16)。
 業務が継続中である場合(S16;NO)、処理はS12に進む。そして、後続するS13~S16の処理が実行される。このようにして、業務が継続中である間、S12~S16の処理が繰り返される。全てのユーザUの業務が終了した場合(S16;YES)、個人仮想環境合成処理は終了する(終了)。
 以上のような個人仮想環境合成処理が、例えば、リアルタイムで実行される業務に対して適用される。
 1.2.2 集中度を算出する処理
 次に、実施形態に係る個人仮想環境合成装置における集中度を算出する処理について説明する。図7は、実施形態に係る個人仮想環境合成装置において検出される、ユーザが業務を中断した時刻を説明するための図である。図7及び以下の説明では、ユーザUが業務を開始してから、n回目に業務を中断したと検出された時刻が、時刻T(fout(n))として示される。
 集中度を算出する処理において、集中度取得部26は、図7に示される各ユーザUが業務を中断した時刻のデータに基づき、下記式(1)で表される集中度の変化率を算出する。
Figure JPOXMLDOC01-appb-M000001
 そして、集中度取得部26は、下記式(2)で表されるシグモイド関数を用いて、集中度の変化率fd(n)を、0~1の値で示される値fs(n)に変換する。
Figure JPOXMLDOC01-appb-M000002
 それから、集中度取得部26は、下記式(3)で表される相乗平均(幾何平均)を用いて、上述の値fs(n)に基づき、時刻T(fout(n))における集中度fr(n)を算出する。
Figure JPOXMLDOC01-appb-M000003
 以上のようにして、各ユーザUの集中度fr(n)(パーセント)が算出される。
 1.2.3 制御比率を算出する処理
 次に、実施形態に係る個人仮想環境合成装置における制御比率を算出する処理について説明する。図8は、実施形態に係る個人仮想環境合成処理における集中度と制御比率との関係の一例を説明するためのグラフである。図8では、例として、集中度が横軸に示され、制御比率が縦軸に示される。
 制御比率を算出する処理において、環境制御部27は、例えば、個人特性情報241に基づいて、各ユーザUについて、集中度と、ウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率の各々との間の制御関数を決定する。
 以下では、ウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率それぞれの制御関数の一例が示される。しかしながら、これらに限られず、ウィンドウサイズ比率、会話音量比率、通知音量比率、及び背景音量比率の各々について、ユーザUごとに異なる制御関数であってもよい。
 環境制御部27は、例えば、下記式(4)を用いて、時刻T(fout(n))における集中度fr(n)に基づき、ウィンドウサイズ比率srw(n)(パーセント)を算出する。図8に示されるように、集中度fr(n)が高いほど、ウィンドウサイズ比率srw(n)は大きくなる。環境制御部27は、例えば、各ユーザUのウィンドウサイズ比率の最適値及び境界値に基づいて、ウィンドウサイズ比率の制御関数を決定することができる。
Figure JPOXMLDOC01-appb-M000004
 また、環境制御部27は、例えば、下記式(5)を用いて、時刻T(fout(n))における集中度fr(n)に基づき、会話音量比率src(n)(パーセント)を算出する。集中度fr(n)が高いほど、会話音量比率src(n)は小さくなる。環境制御部27は、例えば、各ユーザUの会話音量比率の最適値及び境界値に基づいて、会話音量比率の制御関数を決定することができる。
Figure JPOXMLDOC01-appb-M000005
 また、環境制御部27は、例えば、下記式(6)を用いて、時刻T(fout(n))における集中度fr(n)に基づき、通知音量比率srn(n)(パーセント)を算出する。図8に示されるように、集中度fr(n)が高いほど、通知音量比率srn(n)は小さくなる。環境制御部27は、例えば、各ユーザUの会話音データについての通知音量比率の最適値及び境界値に基づいて、通知音量比率の制御関数を決定することができる。
Figure JPOXMLDOC01-appb-M000006
 なお、環境制御部27は、例えば、上記式(6)と同様の式を用いて、背景音量比率srb(n)(パーセント)を算出することができる。なお、背景音量比率についても、環境制御部27は、例えば、各ユーザUの背景音量比率の最適値及び境界値に基づいて、背景音量比率の制御関数を決定することができる。
 1.3 実施形態に係る効果
 実施形態によれば、個人仮想環境合成装置15は、ユーザUの集中度を算出する集中度取得部26と、算出された集中度に基づいて、ユーザUの周辺環境に対する遮蔽率を変更する環境制御部27と、を備える。個人仮想環境合成装置15は、例えば、集中度が高いユーザUの周辺環境に対する遮蔽率を高くするように構成される。また、個人仮想環境合成装置15は、例えば、集中度が低下しているユーザUの周辺環境に対する遮蔽率を低くするように構成される。以上のように、実施形態によれば、個人仮想環境合成装置15は、各ユーザUの集中度に応じて、周辺環境に対する遮蔽率を段階的に変化させることができる。このため、ユーザUの集中力を自然に高めたり、集中度が低下したりした際に、緊張を自然に緩めることができる。したがって、ユーザUの疲労が過度に蓄積することを抑制し、ユーザUが業務を無理なく実施するよう促進することができる。
 また、各ユーザUの制御比率を算出する(遮蔽率を制御する)ための制御関数は、例えば、予め取得される当該ユーザUの特性(個人特性情報241)に基づいて決定される。これにより、実施形態によれば、個人仮想環境合成装置15は、各ユーザUに適した周辺環境の制御を行うことができる。
 2. 変形例等
 なお、上述した実施形態には、種々の変形が適用可能である。
 上述した実施形態では、個人仮想環境合成処理を実行するプログラムが、個人仮想環境合成装置15で実行される場合について説明したが、これに限られない。例えば、個人仮想環境合成処理を実行するプログラムは、クラウド上の計算リソースで実行されてもよい。
 なお、本発明は、上記実施形態に限定されるものではなく、実施段階ではその要旨を逸脱しない範囲で種々に変形することが可能である。また、各実施形態は適宜組み合わせて実施してもよく、その場合組み合わせた効果が得られる。更に、上記実施形態には種々の発明が含まれており、開示される複数の構成要件から選択された組み合わせにより種々の発明が抽出され得る。例えば、実施形態に示される全構成要件からいくつかの構成要件が削除されても、課題が解決でき、効果が得られる場合には、この構成要件が削除された構成が発明として抽出され得る。
 1…個人仮想環境合成システム
 11-1,11-2,11-3,11-4,11-5,11-6…端末
 12…カメラ
 13…マイクロフォンアレイ
 14-1,14-2,14-3,14-4,14-5,14-6…ヘッドフォン
 15…個人仮想環境合成装置
 151…制御回路
 152…通信モジュール
 153…ストレージ
 154…ドライブ
 155…記憶媒体
 21…音声抽出部
 22…会話者推定部
 23…個人特性取得部
 24…記憶部
 25…個人状況取得部
 26…集中度取得部
 27…環境制御部
 28…出力部
 241…個人特性情報
 242…設定情報
 

Claims (8)

  1.  ユーザの集中度を算出する集中度取得部と、
     算出された前記集中度に基づいて、前記ユーザの周辺環境に対する遮蔽率の制御を行う制御部と、
     を備えた、個人仮想環境合成装置。
  2.  前記制御部は、前記遮蔽率の制御の際に、
      前記集中度に基づいて、複数の制御比率を算出し、
      算出された前記複数の制御比率に基づき、前記複数の制御比率のそれぞれに対応する複数の音声データを用いて、前記ユーザに対する出力音声データを合成する、
     ように構成される、
     請求項1記載の個人仮想環境合成装置。
  3.  前記複数の音声データは会話音データ、背景音データ、及び通知音データを含む、
     請求項2記載の個人仮想環境合成装置。
  4.  前記制御部は、前記遮蔽率の制御において、
      前記集中度に応じて前記ユーザが使用する端末を制御する、
     ように構成される、
     請求項1記載の個人仮想環境合成装置。
  5.  前記集中度取得部は、
      前記ユーザの視線、及び前記ユーザが使用する端末の状態のうち少なくともいずれかに基づいて、前記集中度を算出する、
     ように構成される、
     請求項1記載の個人仮想環境合成装置。
  6.  前記制御部は、
      予め取得される前記ユーザの特性に応じて前記遮蔽率の制御を行うように、
     構成される、
     請求項1記載の個人仮想環境合成装置。
  7.  ユーザの集中度を算出することと、
     算出された前記集中度に基づいて、前記ユーザの周辺環境に対する遮蔽率の制御を行うことと、
     を備えた、個人仮想環境合成方法。
  8.  コンピュータを、請求項1乃至請求項6のいずれか1項に記載の個人仮想環境合成装置が備える各部として機能させるためのプログラム。
     
PCT/JP2023/021547 2023-06-09 2023-06-09 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム Ceased WO2024252668A1 (ja)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/JP2023/021547 WO2024252668A1 (ja) 2023-06-09 2023-06-09 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム
JP2025525912A JPWO2024252668A1 (ja) 2023-06-09 2023-06-09

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2023/021547 WO2024252668A1 (ja) 2023-06-09 2023-06-09 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム

Publications (1)

Publication Number Publication Date
WO2024252668A1 true WO2024252668A1 (ja) 2024-12-12

Family

ID=93795215

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2023/021547 Ceased WO2024252668A1 (ja) 2023-06-09 2023-06-09 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム

Country Status (2)

Country Link
JP (1) JPWO2024252668A1 (ja)
WO (1) WO2024252668A1 (ja)

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170256094A1 (en) * 2016-03-01 2017-09-07 International Business Machines Corporation Displaying of augmented reality objects
JP2019152861A (ja) * 2018-03-05 2019-09-12 ハーマン インターナショナル インダストリーズ インコーポレイテッド 集中レベルに基づく、知覚される周囲音の制御
JP2021090136A (ja) * 2019-12-03 2021-06-10 富士フイルムビジネスイノベーション株式会社 情報処理システム及びプログラム
WO2023084945A1 (ja) * 2021-11-10 2023-05-19 株式会社Nttドコモ 出力制御装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20170256094A1 (en) * 2016-03-01 2017-09-07 International Business Machines Corporation Displaying of augmented reality objects
JP2019152861A (ja) * 2018-03-05 2019-09-12 ハーマン インターナショナル インダストリーズ インコーポレイテッド 集中レベルに基づく、知覚される周囲音の制御
JP2021090136A (ja) * 2019-12-03 2021-06-10 富士フイルムビジネスイノベーション株式会社 情報処理システム及びプログラム
WO2023084945A1 (ja) * 2021-11-10 2023-05-19 株式会社Nttドコモ 出力制御装置

Also Published As

Publication number Publication date
JPWO2024252668A1 (ja) 2024-12-12

Similar Documents

Publication Publication Date Title
Chatterjee et al. ClearBuds: wireless binaural earbuds for learning-based speech enhancement
Bentler et al. Digital noise reduction: An overview
US8441515B2 (en) Method and apparatus for minimizing acoustic echo in video conferencing
US9401158B1 (en) Microphone signal fusion
CN103916723B (zh) 一种声音采集方法以及一种电子设备
CN115831155B (zh) 音频信号的处理方法、装置、电子设备及存储介质
CN114255776B (zh) 使用互连电子设备进行音频修改
US11114109B2 (en) Mitigating noise in audio signals
WO2020258328A1 (zh) 一种马达振动方法、装置、系统及可读介质
CN104702787A (zh) 一种应用于移动终端的声音采集方法和移动终端
US20230066600A1 (en) Adaptive noise suppression for virtual meeting/remote education
US20200329322A1 (en) Methods and Apparatus for Auditory Attention Tracking Through Source Modification
CN120981851A (zh) 低时延噪声抑制
CN113906368A (zh) 基于生理观察修改音频
US11818556B2 (en) User satisfaction based microphone array
US20240340605A1 (en) Information processing device and method, and program
JP2024075544A (ja) Vrに応用される環境音ヒアスルー方法、装置、機器および記憶媒体
US11425497B2 (en) Spatial audio zoom
CN115529537B (zh) 一种差分波束形成方法、装置及存储介质
WO2024252668A1 (ja) 個人仮想環境合成装置、個人仮想環境合成方法、及びプログラム
WO2023016032A1 (zh) 一种视频处理方法及电子设备
JP6930280B2 (ja) メディアキャプチャ・処理システム
CN116320144B (zh) 一种音频播放方法及电子设备、可读存储介质
US20200211578A1 (en) Mixed-reality audio intelligibility control
CN111161719B (zh) 一种通过语音操作的ar眼镜及通过语音操作ar眼镜的方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 23940766

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2025525912

Country of ref document: JP

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE