WO2020170946A1 - 音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラム - Google Patents
音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラム Download PDFInfo
- Publication number
- WO2020170946A1 WO2020170946A1 PCT/JP2020/005634 JP2020005634W WO2020170946A1 WO 2020170946 A1 WO2020170946 A1 WO 2020170946A1 JP 2020005634 W JP2020005634 W JP 2020005634W WO 2020170946 A1 WO2020170946 A1 WO 2020170946A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- voice
- data
- audio
- unit
- audio data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0316—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude
- G10L21/0364—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude for improving intelligibility
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M3/00—Automatic or semi-automatic exchanges
- H04M3/42—Systems providing special services or facilities to subscribers
- H04M3/56—Arrangements for connecting several subscribers to a common circuit, i.e. affording conference facilities
- H04M3/568—Arrangements for connecting several subscribers to a common circuit, i.e. affording conference facilities audio processing specific to telephonic conferencing, e.g. spatial distribution, mixing of participants
- H04M3/569—Arrangements for connecting several subscribers to a common circuit, i.e. affording conference facilities audio processing specific to telephonic conferencing, e.g. spatial distribution, mixing of participants using the instant speaker's algorithm
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/40—Processing input control signals of video game devices, e.g. signals generated by the player or derived from the environment
- A63F13/42—Processing input control signals of video game devices, e.g. signals generated by the player or derived from the environment by mapping the input signals into game commands, e.g. mapping the displacement of a stylus on a touch screen to the steering angle of a virtual vehicle
- A63F13/424—Processing input control signals of video game devices, e.g. signals generated by the player or derived from the environment by mapping the input signals into game commands, e.g. mapping the displacement of a stylus on a touch screen to the steering angle of a virtual vehicle involving acoustic input signals, e.g. by using the results of pitch or rhythm extraction or voice recognition
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F13/00—Video games, i.e. games using an electronically generated display having two or more dimensions
- A63F13/85—Providing additional services to players
- A63F13/87—Communicating with other players during game play, e.g. by e-mail or chat
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F13/00—Interconnection of, or transfer of information or other signals between, memories, input/output devices or central processing units
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
- G10L19/04—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis using predictive techniques
- G10L19/16—Vocoder architecture
- G10L19/167—Audio streaming, i.e. formatting and decoding of an encoded audio signal representation into a data stream for transmission or storage purposes
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L21/00—Speech or voice signal processing techniques to produce another audible or non-audible signal, e.g. visual or tactile, in order to modify its quality or its intelligibility
- G10L21/02—Speech enhancement, e.g. noise reduction or echo cancellation
- G10L21/0316—Speech enhancement, e.g. noise reduction or echo cancellation by changing the amplitude
- G10L21/0324—Details of processing therefor
- G10L21/034—Automatic adjustment
-
- A—HUMAN NECESSITIES
- A63—SPORTS; GAMES; AMUSEMENTS
- A63F—CARD, BOARD, OR ROULETTE GAMES; INDOOR GAMES USING SMALL MOVING PLAYING BODIES; VIDEO GAMES; GAMES NOT OTHERWISE PROVIDED FOR
- A63F2300/00—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game
- A63F2300/50—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers
- A63F2300/57—Features of games using an electronically generated display having two or more dimensions, e.g. on a television screen, showing representations related to the game characterized by details of game servers details of game services offered to the player
- A63F2300/572—Communication between players during game play of non game information, e.g. e-mail, chat, file transfer, streaming of audio and streaming of video
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L19/00—Speech or audio signals analysis-synthesis techniques for redundancy reduction, e.g. in vocoders; Coding or decoding of speech or audio signals, using source filter models or psychoacoustic analysis
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L25/00—Speech or voice analysis techniques not restricted to a single one of groups G10L15/00 - G10L21/00
- G10L25/78—Detection of presence or absence of voice signals
-
- H—ELECTRICITY
- H04—ELECTRIC COMMUNICATION TECHNIQUE
- H04M—TELEPHONIC COMMUNICATION
- H04M3/00—Automatic or semi-automatic exchanges
- H04M3/42—Systems providing special services or facilities to subscribers
- H04M3/56—Arrangements for connecting several subscribers to a common circuit, i.e. affording conference facilities
- H04M3/563—User guidance or feature selection
Definitions
- the present invention relates to a voice output control device, a voice output control system, a voice output control method, and a program.
- users may play a game while having a voice chat with another user who is playing a game together with another user who is a remote place such as a viewer of a moving image showing the playing status of the game. Is becoming.
- the present invention has been made in view of the above circumstances, and one of its objects is to provide an audio output control device, an audio output control system, an audio output control method, and a program capable of appropriately thinning out the output of audio data. To provide.
- the audio output control device is a reception unit that receives a plurality of audio data transmitted from mutually different transmitting devices, and an execution result of a voice section detection process for the audio data, Alternatively, based on at least one of the moving averages of the volumes of the voices represented by the voice data, a selection unit for selecting a part of the plurality of voice data and outputting the part of the voice data to be selected. And an output unit.
- the output unit transmits the selected part of the voice data to a reception device capable of voice chat with the transmission device.
- the selection unit may select, from the plurality of audio data, a number of the audio data determined based on the type of the receiving device.
- a decoding unit that decodes the audio data is further included, and the output unit outputs the selected part of the audio data to the decoding unit.
- the selection unit determines, from the plurality of audio data, based on a load of the audio output control device or a communication quality of a computer network to which the audio output control device is connected. Selected number of said audio data.
- An audio output control system includes a plurality of communication devices included in the first group, a plurality of communication devices included in the second group, and a relay device, and the relay devices are mutually Select from a first receiving unit that receives a plurality of audio data transmitted from the communication devices included in the different first group, and a plurality of the audio data that is received by the first receiving unit of the relay device
- a transmitting unit that transmits a part of the audio data to the communication device different from the communication device that has transmitted the audio data, and the communication device included in the second group is a decoding device that decodes the audio data.
- a output unit that outputs a part of the plurality of audio data received by the second receiving unit of the communication device to the decoding unit.
- the communication device that is newly added to the audio output control system is included in the first group or the second group based on the number of communication devices that transmit and receive the audio data to and from each other. Further included is a deciding unit for deciding whether to include in.
- the audio output control method includes a step of receiving a plurality of audio data transmitted from different transmitting devices, an execution result of a voice section detection process on the voice data, or the voice data represents The method includes the steps of selecting a part of the plurality of audio data based on at least one of the moving averages of the sound volume of the sounds, and outputting the selected part of the audio data.
- the program according to the present invention is a procedure for receiving a plurality of voice data transmitted from different transmitting devices, an execution result of a voice section detection process for the voice data, or a volume of a voice represented by the voice data.
- the computer is caused to execute a procedure of selecting a part of the plurality of audio data and a procedure of outputting the selected part of the audio data based on at least one of the moving averages.
- FIG. 3 is a functional block diagram showing an example of functions implemented in the voice chat device and the relay device according to the embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a voice chat device concerning one embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a voice chat device concerning one embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a voice chat device concerning one embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a voice chat device concerning one embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a relay device concerning one embodiment of the present invention. It is a flow figure showing an example of a flow of processing performed in a relay device concerning one embodiment of the present invention.
- FIG. 1 is a diagram showing an example of the overall configuration of a voice chat system 1 according to an embodiment of the present invention.
- the voice chat system 1 includes a voice chat device 10 (10-1, 10-2,..., 10-n), a relay device 12, and a relay device 12, which are mainly configured by computers.
- a management server 14 is included.
- the voice chat device 10, the relay device 12, and the management server 14 are connected to a computer network 16 such as the Internet.
- the voice chat device 10, the relay device 12, and the management server 14 can communicate with each other.
- the management server 14 is, for example, a computer such as a server that manages account information of users who use the voice chat system 1.
- the management server 14 stores, for example, a plurality of account data associated with the user.
- the account data includes, for example, a user ID that is identification information of the user, real name data that indicates the real name of the user, mail address data that indicates the mail address of the user, and the like.
- the voice chat device 10 is a computer capable of inputting/outputting voice of voice chat, such as a game console, a portable game device, a smartphone, and a personal computer.
- the voice chat device 10 includes, for example, a processor 10a, a storage unit 10b, a communication unit 10c, a display unit 10d, an operation unit 10e, a microphone 10f, a speaker 10g, and an encoding/decoding unit 10h. There is.
- the voice chat device 10 may include a camera.
- the processor 10a is a program control device such as a CPU, and executes various types of information processing according to a program stored in the storage unit 10b.
- the storage unit 10b is, for example, a storage element such as a ROM or RAM or a hard disk drive.
- the communication unit 10c is a communication interface for exchanging data with other computers such as the voice chat device 10, the relay device 12, and the management server 14 via the computer network 16, for example.
- the display unit 10d is, for example, a liquid crystal display, and displays a screen generated by the processor 10a, a moving image represented by moving image data received via the communication unit 10c, and the like.
- the operation unit 10e is an operation member for inputting an operation to the processor 10a, for example.
- the microphone 10f is a voice input device used for inputting voice of voice chat, for example.
- the speaker 10g is a voice output device used for outputting voice chat voices, for example.
- the encoding/decoding unit 10h includes, for example, an encoder and a decoder.
- the encoder generates audio data representing the audio by encoding the audio input.
- the decoder also decodes the input audio data and outputs the audio represented by the audio data.
- the relay device 12 is, for example, a computer such as a server that relays the voice data described above.
- the relay device 12 includes, for example, a processor 12a, a storage unit 12b, and a communication unit 12c.
- the processor 12a is, for example, a program control device such as a CPU, and executes various types of information processing according to a program stored in the storage unit 12b.
- the storage unit 12b is, for example, a storage element such as a ROM or RAM or a hard disk drive.
- the communication unit 12c is a communication interface for exchanging data with the computer such as the voice chat device 10 and the management server 14.
- users who use the voice chat system 1 can enjoy voice chat with each other.
- the voice chat may be performed while sharing a moving image representing the play situation of the game being played by some or all of the users participating in the voice chat.
- voice chat by a plurality of users is possible.
- the plurality of users participating in the voice chat belong to a group called a party.
- the user of the voice chat system 1 according to the present embodiment can create a new party or participate in an already created party by performing a predetermined operation.
- FIG. 3 is a diagram schematically showing an example of transmission of voice data in the voice chat system 1 according to the present embodiment.
- FIG. 3 shows an example of transmission of voice data when the voice chat device 10-5 used by the user E receives voice data.
- the voice data transmitted from each of the voice chat devices 10-1 to 10-4 is transmitted to the relay device 12. Then, the relay device 12 selects a part of the plurality of audio data. Then, the relay device 12 transmits the selected part of the voice data to the voice chat device 10-5. As described above, in the example of FIG. 3, the voice data received by the voice chat device 10-5 is thinned out by the relay device 12.
- a part of a plurality of voice data transmitted from each voice chat device 10 different from the voice chat device 10 is partially transmitted. It is selected by the relay device 12. Then, the selected part of the voice data is transmitted to the voice chat device 10.
- FIG. 4 is a diagram schematically showing another example of transmission of voice data in the voice chat system 1 according to the present embodiment.
- FIG. 4 shows an example of transmission of voice data when the voice chat device 10-5 used by the user E receives voice data.
- the voice data transmitted from each of the voice chat devices 10-1 to 10-4 is transmitted to the communication unit 10c of the voice chat device 10-5.
- the processor 10a of the voice chat device 10-5 selects a part of the plurality of voice data.
- the processor 10a of the voice chat device 10-5 outputs the selected part of the voice data to the encoding/decoding unit 10h of the voice chat device 10-5.
- the voice data input to the encoding/decoding unit 10h of the voice chat device 10-5 is thinned out by the processor 10a of the voice chat device 10-5.
- some of the plurality of voice data transmitted from the voice chat device 10 different from the voice chat device 10 are part of the voice chat device 10. Selected by the processor 10a. Then, a part of the selected voice data is output to the encoding/decoding unit 10h of the voice chat device 10.
- FIG 5 and 6 are diagrams schematically showing still another example of transmission of voice data in the voice chat system 1 according to the present embodiment.
- FIGS. 5 and 6 it is assumed that the user A, the user B, the user C, the user D, the user E, and the user F are participating in the party.
- the users A to F use the voice chat devices 10-1 to 10-6, respectively.
- the voice chat devices 10-1 to 10-3 exchange voice data with each other through a P2P connection.
- the voice chat devices 10-4 to 10-6 exchange voice data with other voice chat devices 10 via the relay device 12.
- the voice data transmitted from the voice chat device 10-1 to the voice chat device 10-2 is directly transmitted to the voice chat device 10-2 without passing through the relay device 12.
- the voice data transmitted from the voice chat device 10-1 to the voice chat device 10-3 is directly transmitted to the voice chat device 10-3 without passing through the relay device 12.
- the voice data transmitted from the voice chat device 10-1 to the voice chat devices 10-4 to 10-6 is transmitted to the relay device 12.
- voice data transmitted to the P2P connected voice chat device 10 is directly transmitted, and is transmitted to the voice chat device 10 not connected to the P2P connection.
- the voice data is transmitted to the relay device 12.
- the voice data transmitted from the voice chat device 10 to another voice chat device 10 is transmitted to the relay device 12.
- FIG. 6 shows an example of transmission of voice data when the voice chat device 10-1 used by the user A receives voice data in the situation shown in FIG.
- the voice data transmitted from each of the voice chat devices 10-4 to 10-6 is transmitted to the relay device 12. Then, the relay device 12 selects a part of the plurality of audio data. Then, the relay device 12 transmits the selected part of the voice data to the voice chat device 10-1. As described above, in the example of FIG. 6, the voice data received from the relay device 12 by the voice chat device 10-1 is thinned out by the relay device 12.
- the voice data transmitted from the voice chat devices 10-2 and 10-3 are also transmitted to the communication unit 10c of the voice chat device 10-1.
- the processor 10a of the voice chat device 10-1 selects a part of the plurality of voice data.
- the processor 10a of the voice chat device 10-1 outputs a part of the selected voice data to the encoding/decoding unit 10h of the voice chat device 10-1.
- the voice data input to the encoding/decoding unit 10h of the voice chat device 10-1 is thinned out by the processor 10a of the voice chat device 10-1.
- the relay device 12 selects some of the plurality of voice data transmitted from the voice chat devices 10-4 to 10-6. .. Then, the selected part of the voice data is transmitted to the voice chat device 10. In addition, a part of the plurality of voice data received by the communication unit 10c of the voice chat device 10 is selected by the processor 10a of the voice chat device 10. Then, a part of the selected voice data is output to the encoding/decoding unit 10h of the voice chat device 10.
- the load on the voice chat device 10 and the relay device 12 increases as the number of users participating in the voice chat increases.
- the communication volume of the voice data communication path increases.
- the communication amount of the voice data flowing through the computer network 16 can be suppressed.
- the operation cost of the relay device 12 can be suppressed by adopting the example of FIG.
- the voice data is selected in the relay device 12 and the voice chat device 10 in a distributed manner.
- the voice data is selected in the relay device 12 and the voice chat device 10 in a distributed manner.
- party management data illustrated in FIG. 7, which is stored in the management server 14, for example.
- the party management data includes a party ID, which is identification information of the party, and user data associated with users participating in the party.
- the user data includes a user ID, connection destination address data, type data, a P2P connection flag, and the like.
- the user ID is, for example, identification information of the user.
- the connection destination address data is data indicating the address of the voice chat device 10 used by the user, for example.
- the type data is, for example, data indicating the type of the voice chat device 10 used by the user.
- examples of types of the voice chat device 10 include a game console, a portable game device, a smartphone, a personal computer, and the like, as described above.
- the P2P connection flag is, for example, a flag indicating whether or not the voice chat device 10 used by the user is a voice chat device 10 that performs P2P connection.
- the voice chat device 10 that performs P2P connection refers to the voice chat device 10 that is P2P-connected to the voice chat device 10 that is used by some or all other users who participate in the party. ..
- the voice chat devices 10-1 to 10-3 correspond to the voice chat device 10 that performs P2P connection.
- the voice chat devices 10-1 to 10-5 shown in FIG. 4 also correspond to the voice chat device 10 for P2P connection.
- the voice chat device 10 that does not perform the P2P connection refers to the voice chat device 10 that is connected to the voice chat device 10 used by the user for all other users who participate in the party via the relay device 12. ..
- the voice chat devices 10-4 to 10-6 correspond to the voice chat device 10 that does not make the P2P connection.
- the voice chat devices 10-1 to 10-5 shown in FIG. 3 also correspond to the voice chat device 10 that does not perform P2P connection.
- the value of the P2P connection flag is set to 1 for the user data of the user who uses the voice chat device 10 for P2P connection.
- the value of the P2P connection flag is set to 0.
- FIG. 7 exemplifies party management data having a party ID of 001, which is associated with a party in which six users participate.
- the party management data shown in FIG. 7 includes six pieces of user data associated with users who participate in the party.
- a user whose user ID is aaa, a user who is bbb, a user who is ccc, a user who is ddd, a user who is eee, and a user who is fff are user A, user B, user C, user D, respectively.
- the user A, the user B, the user C, the user D, the user E, and the user F use the voice chat devices 10-1, 10-2, 10-3, 10-4, 10-5, 10-6, respectively. It is assumed that you are using.
- a copy of the party management data stored in the management server 14 is transmitted to the voice chat device 10 and the relay device 12 used by the users who participate in the party associated with the party management data. It Then, a copy of the party management data stored in the management server 14 is stored in the storage unit 10b of the voice chat device 10 and the storage unit 12b of the relay device 12. Therefore, the voice chat device 10 used by the user who participates in the party can specify the address of the voice chat device 10 used by another user who participates in the party.
- the storage unit 10b of the voice chat device 10 also stores data indicating the address of the relay device 12. Therefore, the voice chat device 10 can specify the address of the relay device 12.
- the party management data stored in the management server 14 is updated in accordance with, for example, a user's participation operation in the party. Then, every time the party management data stored in the management server 14 is updated, a copy of the updated party management data is used by the user who participates in the party associated with the party management data. It is also transmitted to the relay device 12. Then, the copy of the party management data stored in the storage unit 10b of the voice chat device 10 and the storage unit 12b of the relay device 12 is updated.
- the latest information shown in the party management data is shared by the voice chat device 10 used by the users who participate in the party associated with the party management data.
- the voice chat device 10 that makes a P2P connection directly sends the voice data that is to be sent to another voice chat device 10 that makes a P2P connection to the voice chat device 10. Then, the voice chat device 10 that makes the P2P connection sends to the relay device 12 the voice data to be sent to the other voice chat device 10 that does not make the P2P connection.
- the voice chat device 10 that does not make the P2P connection transmits to the relay device 12 the voice data to be transmitted to another voice chat device 10.
- a part of the plurality of audio data May be selected.
- the voice chat device 10 generates voice data corresponding to a predetermined period (for example, 20 milliseconds or 40 milliseconds) by encoding the voice input over the period, for each period. You may do it.
- a predetermined period for example, 20 milliseconds or 40 milliseconds
- the voice chat device 10 executes a known voice activity detection (VAD) process on the voice data to determine whether or not the voice data represents a human voice. May be determined. Then, the voice chat device 10 may generate VAD data indicating whether or not the voice data represents a human voice.
- VAD voice activity detection
- the voice chat device 10 may specify the volume of the voice represented by the voice data. Then, the voice chat device 10 may generate volume data indicating the volume of the voice represented by the voice data.
- the voice chat device 10 stores the voice data associated with the identification information of the voice chat device 10 (for example, the user ID corresponding to the voice chat device 10), the VAD data described above, and the volume data described above. You may send it. Further, data indicating a period corresponding to the audio data, such as a time stamp, may be associated with the audio data.
- the relay device 12 which receives a plurality of voice data from different voice chat devices 10, may select a part of the plurality of voice data received during the period, for each predetermined period. ..
- the selection may be performed based on VAD data or volume data associated with the audio data.
- a part of the selected audio data may be transmitted.
- the voice chat device 10 that receives a plurality of voice data from different voice chat devices 10 may select a part of the plurality of voice data received during the period, for each predetermined period. Good.
- the selection may be performed based on VAD data or volume data associated with the audio data.
- a part of the selected voice data may be output to the encoding/decoding unit 10h of the voice chat device 10.
- the functions implemented in the voice chat system 1 according to the present embodiment and the voice chat system 1 according to the present embodiment will be mainly performed with respect to the selection of the voice data and the transmission of the selected voice data. The processing will be further described.
- FIG. 8 is a functional block diagram showing an example of functions implemented in the voice chat device 10 and the relay device 12 according to this embodiment. It should be noted that the voice chat device 10 and the relay device 12 according to the present embodiment do not need to have all of the functions shown in FIG. 8, and may have functions other than those shown in FIG. ..
- the voice chat device 10 functionally, for example, the party management data storage unit 20, the party management unit 22, the voice reception unit 24, the VAD data generation unit 26, the volume data.
- a generation unit 28, a voice data transmission unit 30, a voice data reception unit 32, a selection unit 34, a selected voice data output unit 36, and a voice output unit 38 are included.
- the party management data storage unit 20 is implemented mainly by the storage unit 10b.
- the party management unit 22 is mainly mounted on the processor 10a and the communication unit 10c.
- the voice receiving unit 24 is mainly mounted with the microphone 10f and the encoding/decoding unit 10h.
- the VAD data generation unit 26, the volume data generation unit 28, the selection unit 34, and the selected voice data output unit 36 are mainly mounted on the processor 10a.
- the voice data transmitting unit 30 and the voice data receiving unit 32 are mainly mounted on the communication unit 10c.
- the audio output unit 38 is mainly mounted with the encoding/decoding unit 10h and the speaker 10g.
- the above functions are implemented by executing a program, which is installed in the voice chat device 10 which is a computer, including instructions corresponding to the above functions on the processor 10a.
- This program is supplied to the voice chat device 10 via a computer-readable information storage medium such as an optical disc, a magnetic disc, a magnetic tape, a magneto-optical disc, or a flash memory, or via the Internet or the like.
- the relay device 12 functionally, for example, the party management data storage unit 40, the party management unit 42, the audio data receiving unit 44, the selection unit 46, the audio data.
- a transmitter 48 is included.
- the party management data storage unit 40 is implemented mainly by the storage unit 12b.
- the party management unit 42 is implemented mainly by the processor 12a and the communication unit 12c.
- the voice data receiving unit 44 and the voice data transmitting unit 48 are mainly mounted on the communication unit 12c.
- the selection unit 46 is mainly mounted on the processor 12a.
- the above functions are implemented by executing a program, which is installed in the relay device 12 which is a computer, including a command corresponding to the above functions, on the processor 12a.
- This program is supplied to the relay device 12 via a computer-readable information storage medium such as an optical disc, a magnetic disc, a magnetic tape, a magneto-optical disc, and a flash memory, or via the Internet or the like.
- the party management data storage unit 20 of the voice chat device 10 and the party management data storage unit 40 of the relay device 12 store, for example, the party management data illustrated in FIG. 7 in this embodiment.
- the party management unit 22 of the voice chat device 10 receives the party management data stored in the party management data storage unit 20 in response to the reception of the party management data transmitted from the management server 14. Update to party management data.
- the party management unit 42 of the relay device 12 receives the party management data stored in the party management data storage unit 40 in response to the reception of the party management data transmitted from the management server 14, for example. Update to management data.
- the management server 14 adds the user data including the user ID of the user to the party management data including the party ID of the party.
- the user data will be referred to as additional user data.
- the connection destination address data of the additional user data the address of the voice chat device 10 used by the user is set.
- a value indicating the type of the voice chat device 10 is set in the type data of the additional user data.
- the value of the P2P connection flag of the additional user data is set to 1 or 0 as described above.
- the value of the P2P connection flag of the user data may be determined based on the number of voice chat devices 10 that mutually transmit and receive voice data representing voice of the voice chat.
- the value of the P2P connection flag of the user data may be determined based on the number of user data included in the party management data corresponding to the party.
- the value of the P2P connection flag of the additional user data may be set to 1. ..
- the value of the P2P connection flag of the additional user data may be set to 0.
- voice data can be transmitted and received by P2P connection while the number of voice chat devices 10 that transmit and receive voice data is small.
- the relay device 12 is not used for transmitting/receiving the audio data. Therefore, the load on the relay device 12 can be suppressed while the number of voice chat devices 10 that transmit and receive voice data to each other is small.
- the voice chat device 10 used by the user who performed the participation operation is the voice chat device 10 used by another user of the same party.
- P2P connection may be attempted for each of the above.
- the value of the P2P connection flag of the additional user data may be set to 1.
- the value of the P2P connection flag of the additional user data may be set to 0.
- the additional user data is stored in the voice chat device 10 used by the users participating in the party.
- the party management data is updated.
- the party management data stored in the relay device 12 is also updated.
- the voice receiving unit 24 receives voice of voice chat, for example.
- the voice receiving unit 24 may generate voice data representing the voice by encoding the voice.
- the VAD data generation unit 26 generates the above-mentioned VAD data based on the voice data generated by the voice reception unit 24, for example.
- the volume data generation unit 28 generates the above-described volume data based on the voice data generated by the voice reception unit 24, for example.
- the voice data transmitting unit 30 of the voice chat device 10 transmits, for example, voice data representing a voice received by the voice receiving unit 24.
- the voice data transmitting unit 30 may transmit voice data associated with the identification information of the voice chat device 10.
- the voice data transmission unit 30 outputs the voice data in which the identification information of the voice chat device 10, the VAD data generated by the VAD data generation unit 26, and the volume data generated by the volume data generation unit 28 are associated with each other. You may send it.
- data indicating a period corresponding to the audio data such as a time stamp, may be associated with the audio data.
- the voice data transmitting unit 30 of the voice chat device 10 transmits voice data to the voice chat device 10 used by another user who is participating in the same party as the user who uses the voice chat device 10. You may. Further, the voice data transmitting unit 30 of the voice chat device 10 relays voice data addressed to the voice chat device 10 used by another user who is participating in the same party as the user who uses the voice chat device 10. May be sent to.
- the voice data receiving unit 32 of the voice chat device 10 receives, for example, voice data in this embodiment.
- the audio data receiving unit 32 may receive a plurality of audio data transmitted from different transmitting devices.
- the voice chat device 10 used by another user who is participating in the same party as the user who uses the voice chat device 10 or the relay device 12 corresponds to the transmission device.
- the voice data receiving unit 32 of the voice chat device 10 may directly transmit voice data from the voice chat device 10 used by another user who is participating in the same party as the user who uses the voice chat device 10. Further, the voice data receiving unit 32 of the voice chat device 10 transmits via the relay device 12 from the voice chat device 10 used by another user who is participating in the same party as the user who uses the voice chat device 10. The received audio data may be received.
- the selection unit 34 of the voice chat device 10 selects a part of the plurality of voice data received by the voice data reception unit 32 of the voice chat device 10.
- the selection unit 34 selects a part of the plurality of audio data based on at least one of the execution result of the audio section detection process on the audio data or the volume of the audio represented by the audio data. You may.
- the selected voice data output unit 36 outputs, for example, a part of the voice data selected by the selection unit 34 of the voice chat device 10.
- a part of the voice data selected by the selection unit 34 of the voice chat device 10 is output to the voice output unit 38.
- the audio output unit 38 decodes the audio data output from the selected audio data output unit 36, for example. Then, in the present embodiment, the audio output unit 38 outputs the audio representing the audio data, which is generated by decoding the audio data, for example.
- the voice data receiving unit 44 of the relay device 12 receives, for example, a plurality of voice data transmitted from different transmitting devices.
- the voice chat device 10 that transmits voice data to the relay device 12 corresponds to the transmission device.
- the voice data receiving unit 44 of the relay device 12 receives, for example, voice data transmitted by the voice data transmitting unit 30 of the voice chat device 10.
- the selection unit 46 of the relay device 12 selects a part of a plurality of audio data received by the audio data reception unit 44 of the relay device 12.
- the selection unit 46 selects a part of the plurality of audio data based on at least one of the execution result of the audio section detection process on the audio data or the volume of the audio represented by the audio data. You may.
- the voice data transmitting unit 48 of the relay device 12 is a part of the voice data received by the voice data receiving unit 44 of the relay device 12, for example, a part of the voice data selected by the selecting unit 46 of the relay device 12.
- the data is transmitted to the receiving device that can perform voice chat with the transmitting device that is the transmission source.
- the voice chat device 10 corresponds to the receiving device.
- the voice data transmitting unit 48 may specify the party to which the user represented by the user ID belongs, based on the user ID associated with the voice data. Then, the voice data transmitting unit 48 transmits the voice data to the voice chat devices 10 used by the users participating in the party other than the voice chat device 10 used by the user associated with the user ID. Good.
- the relay device 12 may receive a plurality of audio data respectively transmitted from the communication devices included in the first group different from each other. Then, the relay device 12 may transmit a part selected from the plurality of voice data to a communication device different from the communication device that transmitted the voice data.
- the voice chat devices 10-4 to 10-6 correspond to the communication devices included in the first group. That is, the voice chat device 10 associated with the user data whose P2P connection flag value is 0 corresponds to the communication device included in the first group.
- the voice chat device 10 included in the second group transmits the above-mentioned part of the voice data transmitted from the relay device 12 and at least one transmitted from the other communication devices included in the second group different from each other.
- One voice data may be received.
- the voice chat device 10 may output a part of the plurality of received voice data to the voice output unit 38 of the voice chat device 10.
- the voice chat device 10-1 corresponds to the voice chat device 10 that receives voice data.
- the voice chat devices 10-1 to 10-3 correspond to the communication devices included in the second group. That is, the voice chat device 10 associated with the user data in which the value of the P2P connection flag is 1 corresponds to the communication device included in the second group.
- the management server 14 includes a communication device newly added to the voice chat system 1 in the first group or in the second group based on the number of communication devices that transmit and receive voice data to and from each other. May be determined. For example, as described above, the management server 14 determines the value of the P2P connection flag included in the user data corresponding to the newly added voice chat device 10, based on the number of voice chat devices 10 that exchange voice data with each other. May be determined.
- FIG. 9 a flowchart of a voice data transmission process performed in the voice chat device 10 according to the present embodiment will be described with reference to a flowchart illustrated in FIG. 9.
- the processing shown in S101 to S107 shown in FIG. 9 is repeatedly executed every predetermined period (for example, every 20 milliseconds or every 40 milliseconds).
- the voice receiving unit 24 generates voice data by encoding the voice received during the period of this loop (S101).
- the VAD data generation unit 26 executes the VAD process on the voice data generated by the process shown in S101, thereby determining whether or not the voice data represents a human voice ( S102).
- the VAD data generation unit 26 generates VAD data according to the determination result in the process shown in S102 (S103).
- the volume data generation unit 28 identifies the volume of the voice represented by the voice data generated by the process shown in S101 (S104).
- volume data generation unit 28 generates volume data indicating the volume specified in the process shown in S104 (S105).
- the voice data transmitting unit 30 identifies the communication device as the transmission destination of the voice data based on the party management data stored in the party management data storage unit 20 (S106).
- the address of the voice chat device 10 that is the destination, the necessity of transmitting the voice data to the relay device 12, and the like are specified.
- the voice data transmitting unit 30 transmits the voice data generated in the process shown in S101 to the destination specified in the process shown in S106 (S107), and returns to the process shown in S101.
- the voice data is associated with the user ID of the user who uses the voice chat device 10, the VAD data generated in the process of S103, and the volume data generated in the process of S105. ing. Further, data indicating the period of this loop, such as a time stamp, may be associated with the audio data.
- the process proceeds to S102 for a predetermined time (for example, 1 second) from the determined timing.
- the processing shown may not be executed.
- VAD data having a value of 1 may be generated in the process shown in S103.
- FIGS. 10A and 10B An example of the flow of voice output processing performed in the voice chat device 10 according to the present embodiment will be described with reference to the flow charts illustrated in FIGS. 10A and 10B.
- the processing shown in S201 to S217 shown in FIGS. 10A and 10B is repeatedly executed every predetermined period (for example, every 20 milliseconds or every 40 milliseconds).
- predetermined period for example, every 20 milliseconds or every 40 milliseconds.
- the maximum number of voice data that can be input by the voice output unit 38 in the voice chat device 10 in one period is predetermined.
- the maximum number is expressed as n1.
- the selection unit 34 identifies the audio data received by the audio data receiving unit 32 during the period of this loop (S201).
- the selection unit 34 confirms whether or not the number of audio data identified by the processing shown in S201 (hereinafter, expressed as m1) is n1 or less (S202).
- m1 is confirmed to be n1 or less in the process shown in S202 (S202: Y).
- the selected audio data output unit 36 outputs all the audio data specified in the process of S201 to the audio output unit 38 (S203).
- the selection unit 34 confirms whether or not the number of pieces of audio data identified by the processing shown in S204 (hereinafter referred to as m2) is n1 or more (S205).
- the selection unit 34 specifies n1 audio data from the plurality of audio data specified in the process of S204, in order from the one having the highest volume indicated by the associated volume data (S206). ..
- n1 audio data is specified in the process shown in S204, in order from the one having the highest volume indicated by the associated volume data (S206). ..
- m2 is the same as n1, in the process shown in S206, all the audio data specified in the process shown in S204 will be specified.
- the selected voice data output unit 36 outputs the n1 voice data specified in the process of S206 to the voice output unit 38 (S207).
- the selecting unit 34 identifies the voice data whose associated VAD data has a value of 0, from the plurality of voice data identified by the process of S201 (S208).
- the selection unit 34 identifies the moving average of the volume of the audio represented by the audio data (S209).
- the moving average of the volume of the voice represented by the voice data is, for example, received from the voice chat device 10 that is the transmission source of the voice data at the latest predetermined number of times or the latest predetermined time (for example, the latest one second). It refers to the average volume of the voice represented by the voice data.
- the selection unit 34 may store at least the audio data received for the latest predetermined times or for the latest predetermined time.
- the selection unit 34 identifies, from among the plurality of audio data identified by the process of S208, the audio data whose moving average identified by the process of S209 is equal to or more than a predetermined threshold value (for example, ⁇ 40 dBOV or more). Yes (S210).
- a predetermined threshold value for example, ⁇ 40 dBOV or more
- the selection unit 34 compares n1 with the total (hereinafter, expressed as m3) of the number of audio data identified by the process shown in S204 and the number of audio data identified by the process shown in S210. Yes (S211).
- the selection unit 34 specifies n1 audio data from the plurality of audio data specified in the process of S204 or S210, in descending order of the volume indicated by the associated volume data ( S212). Then, the selected audio data output unit 36 outputs the n1 audio data specified in the process of S212 to the audio output unit 38 (S213).
- the selection unit 34 outputs a total of n1 pieces of audio data specified in the process of S204 or S210 to the audio output unit 38 (S216).
- the audio output unit 38 decodes the audio data output by the processing shown in S203, S207, S213, S215, or S216, and outputs the audio represented by the audio data (S217). Then, the process returns to S201.
- the selection unit 34 of the voice chat device 10 accurately selects the voice data that is likely to represent the voice of the voice chat from the plurality of voice data. Will be done.
- the audio data selection processing is not limited to the processing shown in the processing examples shown in FIGS. 10A and 10B.
- the selection unit 34 may specify audio data whose associated VAD data has a value of 1 or whose moving average of sound volume is equal to or higher than a predetermined threshold value. Then, the selection unit 34 may specify n1 pieces of audio data in order from the audio data having the highest volume indicated by the associated volume data.
- FIGS. 11A and 11B An example of the flow of voice data relay processing performed in the relay device 12 according to the present embodiment will be described with reference to the flow charts illustrated in FIGS. 11A and 11B.
- the processing shown in S301 to S316 shown in FIGS. 11A and 11B is repeatedly executed every predetermined period (for example, every 20 milliseconds or every 40 milliseconds). Further, in the following description, it is assumed that the maximum number of audio data that can be transmitted by the relay device 12 in one period is predetermined. Hereinafter, the maximum number is expressed as n2.
- the selection unit 46 of the relay device 12 specifies in advance the plurality of voice chat devices 10 that are the destinations of the voice data, based on the party management data stored in the party management data storage unit 40. To do.
- the selecting unit 46 identifies the audio data received by the audio data receiving unit 44 during the period of this loop (S301).
- the selection unit 46 confirms whether or not the number of audio data identified by the processing shown in S301 (hereinafter, expressed as m4) is n2 or less (S302).
- the voice data transmitting unit 48 transmits all the voice data specified in the process of S301 to the voice chat device 10 that is the destination (S303), and returns to the process of S301.
- the selection unit 46 identifies the voice data whose associated VAD data has a value of 1 from the plurality of voice data identified in the process of S301 (S304).
- the selection unit 46 confirms whether or not the number of pieces of audio data specified in the process shown in S304 (hereinafter referred to as m5) is n2 or more (S305).
- m5 is confirmed to be n2 or more in the process shown in S305 (S305: Y).
- the selection unit 46 specifies n2 pieces of audio data from the plurality of pieces of audio data specified by the processing shown in S204, in descending order of the volume indicated by the associated volume data (S306). ..
- m5 is the same as n2, in the process shown in S306, all the audio data specified in the process shown in S304 will be specified.
- the voice data transmitting unit 48 transmits the n2 voice data specified in the process of S306 to the voice chat device 10 that is the transmission destination (S307), and returns to the process of S301.
- the selection unit 46 identifies the voice data whose associated VAD data has a value of 0 from the plurality of voice data identified by the process shown in S301 (S308).
- the selection unit 46 specifies the moving average of the sound volume of the audio represented by the audio data for each of the audio data identified by the processing shown in S308 (S309).
- the moving average of the volume of the voice represented by the voice data is, for example, received from the voice chat device 10 that is the transmission source of the voice data at the latest predetermined number of times or the latest predetermined time (for example, the latest one second). It refers to the average volume of the voice represented by the voice data.
- the selection unit 46 may store at least the audio data received for the latest predetermined times or for the latest predetermined time.
- the selection unit 46 identifies, from among the plurality of audio data identified by the process of S308, the audio data whose moving average identified by the process of S309 is equal to or more than a predetermined threshold (for example, -40 dBOV or more). Yes (S310).
- a predetermined threshold for example, -40 dBOV or more
- the selection unit 46 compares n2 with the total (hereinafter referred to as m6) of the number of audio data identified by the process shown in S304 and the number of audio data identified by the process shown in S210. Yes (S311).
- the selection unit 46 specifies n2 audio data from the plurality of audio data specified by the processing shown in S304 or S310, in order from the one with the highest volume indicated by the associated volume data ( S312). Then, the voice data transmitting unit 48 transmits the n2 voice data specified in the process of S312 to the voice chat device 10 that is the destination (S313), and returns to the process of S301.
- the selection unit 46 selects (n2-m6) voices from the remaining voice data not specified in the process of S304 or S310, in descending order of the volume indicated by the associated volume data. The data is specified (S314). Then, the selected voice data output unit 36 transmits the total n2 voice data identified by the processing in S304, S310, or S314 to the voice chat device 10 that is the destination (S315), and then in S301. Return to the processing shown.
- the selection unit 46 transmits a total of n2 pieces of voice data specified in the process of S304 or S310 to the voice chat device 10 that is the destination (S316), and returns to the process of S301.
- the selection unit 46 of the relay device 12 accurately selects the voice data that is highly likely to represent the voice of the voice chat from the plurality of voice data.
- the Rukoto is the processing example illustrated in FIGS. 11A and 11B.
- voice data based on VAD data it is possible to reduce the possibility that voice data that represents a sound other than the human voice, such as the sound of hitting a desk or the sound of an ambulance, will be selected. Further, by selecting the voice data based on the moving average of the sound volume, for example, the voice data that does not have been selected by the voice data selection based on the VAD data but may actually be the voice data representing a human voice may be selected. Can be raised.
- the selection unit 34 or the selection unit 46 may select a part of the plurality of audio data based on the moving average of the volume of the audio represented by the audio data. In this way, the audio data can be stably selected.
- the selection unit 46 selects, from among the plurality of voice data, the number of voice data determined based on the type of the voice chat device 10 that is the transmission destination of the voice data. You may. Then, the voice data transmitting unit 48 may transmit the part of the voice data to the voice chat device 10. In this case, the number n2 described above varies depending on the type of the voice chat device 10 that is the transmission destination.
- the voice chat device 10 that is a smartphone transmits a smaller number of voice data than the voice chat device 10 that is a game console.
- the number of audio data determined based on the load of the device or the communication quality of the computer network 16 may be selected.
- the selection unit 34 of the voice chat device 10 may determine the above value n1 based on the load of the voice chat device 10.
- the selection unit 46 of the relay device 12 may determine the above-mentioned value n2 based on the load of the relay device 12 or the communication quality (for example, communication amount) of the computer network 16.
- the voice chat system 1 it is possible to appropriately thin out the output of the voice data representing the voice of the voice chat.
- the transmission output of voice data by the relay device 12 and the output of voice data from the processor 10a of the voice chat device 10 to the encoding/decoding unit 10h can be appropriately thinned.
- the present invention is not limited to the above embodiment.
- the selection unit 34 may store a list of user IDs that represent the users who have recently spoken.
- the list may include n1 user IDs.
- the selecting unit 34 may select a part of the voice data transmitted by the voice chat device 10 used by the user whose user ID is not included in the list.
- the number n3 of audio data selected is smaller than n1.
- the selection unit 34 may delete the n3 user IDs from the list. For example, n3 user IDs may be deleted from the list in order from the earliest timing added to the list. In addition, n3 user IDs may be deleted from the list in order from the smallest volume of the voice represented by the latest voice data.
- the selection unit 34 may add the user IDs associated with the selected n3 audio data to the list.
- the selection unit 34 may select the voice data associated with the user ID included in the list. In this way, the voice data can be stably selected in the voice chat device 10.
- the selection unit 46 may store a list including n2 user IDs. Then, similarly to the selecting unit 34 in the above example, the voice data associated with the user ID included in the list may be selected. By doing so, the selection of the voice data in the relay device 12 can be stably performed.
- the relay device 12 may determine whether the voice data represents a human voice or specify the volume of the voice represented by the voice data. Further, for example, in the voice chat device 10 that receives voice data, it may be determined whether or not the voice data represents a human voice, and the volume of the voice represented by the voice data may be specified.
Landscapes
- Engineering & Computer Science (AREA)
- Multimedia (AREA)
- Physics & Mathematics (AREA)
- Acoustics & Sound (AREA)
- Human Computer Interaction (AREA)
- Signal Processing (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Quality & Reliability (AREA)
- Theoretical Computer Science (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Information Transfer Between Computers (AREA)
- Preparation Of Compounds By Using Micro-Organisms (AREA)
- Micro-Organisms Or Cultivation Processes Thereof (AREA)
- Peptides Or Proteins (AREA)
- Telephonic Communication Services (AREA)
Abstract
音声データの出力を適切に間引くことができる音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラムを提供する。音声データ受信部(44)は、それぞれ互いに異なる送信装置から送信される複数の音声データを受信する。選択部(46)は、音声データに対する音声区間検出処理の実行結果、又は、音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、複数の音声データのうちの一部を選択する。音声データ送信部(48)は、選択される一部の音声データを出力する。
Description
本発明は、音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラムに関する。
近年、ともにゲームをプレイしているユーザやゲームのプレイ状況を表す動画像の閲覧者などといった離れた場所にいる他のユーザとボイスチャットをしながら、ユーザがゲームをプレイすることが行われるようになってきている。
ボイスチャットに参加するユーザの増加に伴って、各ユーザがボイスチャットにおける音声の入出力に利用する装置や音声データを中継するサーバの負荷が高くなる。また、音声データの通信経路の通信量が増える。このことはボイスチャットのサービス品質の低下や、サーバの運用コストや通信コストの増加などを引き起こす。
そのため、多くのユーザが参加するボイスチャットにおいては音声データの出力を適切に間引いて上記の負荷や通信量を適切に抑える必要がある。しかし従来のボイスチャットの技術ではこのようなことが行われていなかった。
本発明は上記実情に鑑みてなされたものであって、その目的の一つは、音声データの出力を適切に間引くことができる音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラムを提供することにある。
上記課題を解決するために、本発明に係る音声出力制御装置は、それぞれ互いに異なる送信装置から送信される複数の音声データを受信する受信部と、前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択する選択部と、選択される前記一部の前記音声データを出力する出力部と、を含む。
また、本発明の一態様では、前記出力部は、選択される前記一部の前記音声データを、前記送信装置とボイスチャットが可能な受信装置に送信する。
この態様では、前記選択部は、前記複数の音声データのうちから、前記受信装置の種類に基づいて決定される数の前記音声データを選択してもよい。
また、本発明の一態様では、前記音声データをデコードするデコード部、をさらに含み、前記出力部は、選択される前記一部の前記音声データを前記デコード部に出力する。
また、本発明の一態様では、前記選択部は、前記複数の音声データのうちから、前記音声出力制御装置の負荷又は前記音声出力制御装置が接続されているコンピュータネットワークの通信品質に基づいて決定される数の前記音声データを選択する。
また、本発明に係る音声出力制御システムは、第1群に含まれる複数の通信装置と、第2群に含まれる複数の通信装置と、中継装置と、を含み、前記中継装置は、それぞれ互いに異なる前記第1群に含まれる前記通信装置から送信される複数の音声データを受信する第1受信部と、前記中継装置の前記第1受信部が受信する前記複数の前記音声データのうちから選択される一部を、当該音声データを送信した前記通信装置とは異なる前記通信装置に送信する送信部と、を含み、前記第2群に含まれる前記通信装置は、前記音声データをデコードするデコード部と、前記中継装置から送信される前記一部の前記音声データ、及び、それぞれ互いに異なる前記第2群に含まれる他の前記通信装置から送信される少なくとも1つの音声データを受信する第2受信部と、前記通信装置の前記第2受信部が受信する複数の前記音声データのうちの一部を前記デコード部に出力する出力部と、を含む。
本発明の一態様では、前記音声データが互いに送受信される前記通信装置の数に基づいて、前記音声出力制御システムに新たに追加される前記通信装置を前記第1群に含めるか前記第2群に含めるかを決定する決定部、をさらに含む。
また、本発明に係る音声出力制御方法は、それぞれ互いに異なる送信装置から送信される複数の音声データを受信するステップと、前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択するステップと、選択される前記一部の前記音声データを出力するステップと、を含む。
また、本発明に係るプログラムは、それぞれ互いに異なる送信装置から送信される複数の音声データを受信する手順、前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択する手順、選択される前記一部の前記音声データを出力する手順、をコンピュータに実行させる。
図1は、本発明の一実施形態に係るボイスチャットシステム1の全体構成の一例を示す図である。図1に示すように、ボイスチャットシステム1には、いずれもコンピュータを中心に構成された、ボイスチャット装置10(10-1、10-2、・・・、10-n)、中継装置12、管理サーバ14が、含まれている。ボイスチャット装置10、中継装置12、管理サーバ14は、インターネットなどのコンピュータネットワーク16に接続されている。ボイスチャット装置10、中継装置12、管理サーバ14は、互いに通信可能となっている。
管理サーバ14は、例えば、ボイスチャットシステム1を利用するユーザのアカウント情報などを管理するサーバ等のコンピュータである。管理サーバ14は、例えば、ユーザに対応付けられるアカウントデータを複数記憶する。ここでアカウントデータには、例えば、当該ユーザの識別情報であるユーザID、当該ユーザの実名を示す実名データ、当該ユーザのメールアドレスを示すメールアドレスデータ、などが含まれる。
ボイスチャット装置10は、例えば、ゲームコンソール、携帯型ゲーム装置、スマートフォン、パーソナルコンピュータなどといった、ボイスチャットの音声の入出力が可能なコンピュータである。
図2Aに示すように、ボイスチャット装置10には、例えば、プロセッサ10a、記憶部10b、通信部10c、表示部10d、操作部10e、マイク10f、スピーカ10g、エンコード・デコード部10hが含まれている。なお、ボイスチャット装置10がカメラを備えていてもよい。
プロセッサ10aは、例えばCPU等のプログラム制御デバイスであって、記憶部10bに記憶されたプログラムに従って各種の情報処理を実行する。
記憶部10bは、例えばROMやRAM等の記憶素子やハードディスクドライブなどである。
通信部10cは、例えばコンピュータネットワーク16を介して、他のボイスチャット装置10、中継装置12、管理サーバ14などといったコンピュータとの間でデータを授受するための通信インタフェースである。
表示部10dは、例えば液晶ディスプレイ等であり、プロセッサ10aが生成する画面や、通信部10cを介して受信する動画像データが表す動画像などを表示させる。
操作部10eは、例えばプロセッサ10aに対する操作入力を行うための操作部材である。
マイク10fは、例えばボイスチャットの音声の入力に用いられる音声入力デバイスである。
スピーカ10gは、例えばボイスチャットの音声の出力に用いられる音声出力デバイスである。
エンコード・デコード部10hは、例えばエンコーダとデコーダとを含む。当該エンコーダは、入力される音声をエンコードすることにより当該音声を表す音声データを生成する。また、当該デコーダは、入力される音声データをデコードして、当該音声データが表す音声を出力する。
中継装置12は、本実施形態では例えば、上述の音声データを中継するサーバ等のコンピュータである。
図2Bに示すように、本実施形態に係る中継装置12には、例えば、プロセッサ12a、記憶部12b、通信部12c、が含まれている。
プロセッサ12aは、例えばCPU等のプログラム制御デバイスであって、記憶部12bに記憶されたプログラムに従って各種の情報処理を実行する。
記憶部12bは、例えばROMやRAM等の記憶素子やハードディスクドライブなどである。
通信部12cは、ボイスチャット装置10や、管理サーバ14等のコンピュータとの間でデータを授受するための通信インタフェースである。
本実施形態では、ボイスチャットシステム1を利用するユーザ同士が互いにボイスチャットを楽しめるようになっている。ここで例えば、ボイスチャットに参加している一部又は全部のユーザがプレイ中であるゲームのプレイ状況を表す動画像を共有しながらボイスチャットが行われるようにしてもよい。
また本実施形態では、複数のユーザによるボイスチャットが可能である。ここで本実施形態では、ボイスチャットに参加している複数のユーザは、パーティと呼ばれるグループに属することとする。本実施形態に係るボイスチャットシステム1のユーザは、所定の操作を行うことで、新規のパーティの作成や既に作成されているパーティへの参加を行うことができる。
図3は、本実施形態に係るボイスチャットシステム1における音声データの送信の一例を模式的に示す図である。
図3の例では、パーティに、ユーザA、ユーザB、ユーザC、ユーザD、ユーザEが参加していることとする。そして、ユーザA~ユーザEは、それぞれ、ボイスチャット装置10-1~10-5を利用することとする。そして図3には、ユーザEが利用するボイスチャット装置10-5が音声データを受信する際の音声データの送信の一例が示されている。
図3の例では、ボイスチャット装置10-1~10-4のそれぞれから送信される音声データは、中継装置12に送信される。そして中継装置12は、これら複数の音声データのうちの一部を選択する。そして中継装置12は、選択された一部の音声データをボイスチャット装置10-5に送信する。このように図3の例では、ボイスチャット装置10-5が受信する音声データは中継装置12によって間引かれていることとなる。
なお図3の例では、ボイスチャット装置10-1~10-4についても同様にして、当該ボイスチャット装置10とは異なるボイスチャット装置10からそれぞれ送信される複数の音声データのうちの一部が中継装置12によって選択される。そして、選択された一部の音声データが当該ボイスチャット装置10に送信される。
図4は、本実施形態に係るボイスチャットシステム1における音声データの送信の別の一例を模式的に示す図である。
図4の例でも、パーティに、ユーザA、ユーザB、ユーザC、ユーザD、ユーザEが参加していることとする。そして、ユーザA~ユーザEは、それぞれ、ボイスチャット装置10-1~10-5を利用することとする。そして図4には、ユーザEが利用するボイスチャット装置10-5が音声データを受信する際の音声データの送信の一例が示されている。
図4の例では、ボイスチャット装置10-1~10-4のそれぞれから送信される音声データは、ボイスチャット装置10-5の通信部10cに送信される。そしてボイスチャット装置10-5のプロセッサ10aは、これら複数の音声データのうちの一部を選択する。そしてボイスチャット装置10-5のプロセッサ10aは、選択された一部の音声データをボイスチャット装置10-5のエンコード・デコード部10hに出力する。このように図4の例では、ボイスチャット装置10-5のエンコード・デコード部10hに入力される音声データはボイスチャット装置10-5のプロセッサ10aによって間引かれていることとなる。
なお、ボイスチャット装置10-1~10-4についても同様にして、当該ボイスチャット装置10とは異なるボイスチャット装置10からそれぞれ送信される複数の音声データのうちの一部が当該ボイスチャット装置10のプロセッサ10aによって選択される。そして、選択された一部の音声データが当該ボイスチャット装置10のエンコード・デコード部10hに出力される。
図5及び図6は、本実施形態に係るボイスチャットシステム1における音声データの送信のさらに別の一例を模式的に示す図である。
図5及び図6の例では、パーティに、ユーザA、ユーザB、ユーザC、ユーザD、ユーザE、ユーザFが参加していることとする。そして、ユーザA~ユーザFは、それぞれ、ボイスチャット装置10-1~10-6を利用することとする。
図5に示すように、ここでは例えば、ボイスチャット装置10-1~10-3はP2P接続により、互いに音声データが送受信される。一方で、ボイスチャット装置10-4~10-6は、中継装置12を経由して、他のボイスチャット装置10との間での音声データが送受信される。
そのため例えば、ボイスチャット装置10-1がボイスチャット装置10-2に宛てて送信する音声データは、中継装置12を経由せずにボイスチャット装置10-2に直接送信される。また、ボイスチャット装置10-1がボイスチャット装置10-3に宛てて送信する音声データは、中継装置12を経由せずにボイスチャット装置10-3に直接送信される。
一方で、ボイスチャット装置10-1がボイスチャット装置10-4~10-6に宛てて送信する音声データは、中継装置12に送信される。
ボイスチャット装置10-2、10-3についても同様に、P2P接続されたボイスチャット装置10に宛てて送信される音声データは直接送信され、P2P接続されていないボイスチャット装置10に宛てて送信される音声データは中継装置12に送信される。
また、ボイスチャット装置10-4~10-6については、当該ボイスチャット装置10が他のボイスチャット装置10に宛てて送信する音声データは、中継装置12に送信される。
そして図6には、図5に示す状況においてユーザAが利用するボイスチャット装置10-1が音声データを受信する際の音声データの送信の一例が示されている。
図6の例では、上述のように、ボイスチャット装置10-4~10-6のそれぞれから送信される音声データは、中継装置12に送信される。そして中継装置12は、これら複数の音声データのうちの一部を選択する。そして中継装置12は、選択された一部の音声データをボイスチャット装置10-1に送信する。このように図6の例では、ボイスチャット装置10-1が中継装置12から受信する音声データは中継装置12によって間引かれていることとなる。
そして図6の例では、中継装置12から送信される音声データに加え、ボイスチャット装置10-2、10-3から送信される音声データについても、ボイスチャット装置10-1の通信部10cに送信される。そしてボイスチャット装置10-1のプロセッサ10aは、これら複数の音声データのうちの一部を選択する。そしてボイスチャット装置10-1のプロセッサ10aは、選択された一部の音声データをボイスチャット装置10-1のエンコード・デコード部10hに出力する。このように図6の例では、ボイスチャット装置10-1のエンコード・デコード部10hに入力される音声データはボイスチャット装置10-1のプロセッサ10aによって間引かれていることとなる。
なお、ボイスチャット装置10-2、10-3についても同様にして、ボイスチャット装置10-4~10-6からそれぞれ送信される複数の音声データのうちの一部が中継装置12によって選択される。そして、選択された一部の音声データが当該ボイスチャット装置10に送信される。また、当該ボイスチャット装置10の通信部10cが受信する複数の音声データのうちの一部が当該ボイスチャット装置10のプロセッサ10aによって選択される。そして、選択された一部の音声データが当該ボイスチャット装置10のエンコード・デコード部10hに出力される。
ボイスチャットに参加するユーザの増加に伴って、ボイスチャット装置10や中継装置12の負荷が高くなる。また、音声データの通信経路の通信量が増える。図3の例によれば、中継装置12からボイスチャット装置10-5に送信される音声データが中継装置12によって間引かれているため、コンピュータネットワーク16を流れる音声データの通信量(ネットワークトラフィック)を抑えることができる。また例えば、中継装置12が、データの送信量が増えるに従って運用コストが増えるようなものである場合は、図3の例のようにすることで、中継装置12の運用コストを抑えることができる。
また図4の例によれば、ボイスチャット装置10-5のエンコード・デコード部10hに入力される音声データが間引かれているため、ボイスチャット装置10-5におけるデコードの処理負荷を抑えることができる。
また図5及び図6の例によれば、音声データの選択が中継装置12とボイスチャット装置10とで分散して行われることとなる。その結果、コンピュータネットワーク16を流れる音声データの通信量の抑制、中継装置12の運用コストの抑制、及び、ボイスチャット装置10-1におけるデコードの処理負荷の抑制がバランスよく達成できる。
本実施形態では、図5に示すような音声データの通信経路などといった、パーティに関する情報は、例えば管理サーバ14に記憶される、図7に例示するパーティ管理データによって管理されている。図7に示すように、パーティ管理データには、パーティの識別情報であるパーティIDと、それぞれ当該パーティに参加しているユーザに対応付けられるユーザデータと、が含まれる。そして、ユーザデータには、ユーザID、接続先アドレスデータ、種類データ、P2P接続フラグ、などが含まれている。
ユーザIDは、例えば当該ユーザの識別情報である。接続先アドレスデータは、例えば当該ユーザが利用するボイスチャット装置10のアドレスを示すデータである。種類データは、例えば当該ユーザが利用するボイスチャット装置10の種類を示すデータである。ここでボイスチャット装置10の種類の例としては、上述のように、ゲームコンソール、携帯型ゲーム装置、スマートフォン、パーソナルコンピュータなどが挙げられる。P2P接続フラグは、例えば当該ユーザが利用するボイスチャット装置10がP2P接続を行うボイスチャット装置10であるか否かを示すフラグである。
ここで例えば、P2P接続を行うボイスチャット装置10とは、パーティに参加する一部又は全部の他のユーザについて、当該ユーザが利用するボイスチャット装置10とP2P接続がされるボイスチャット装置10を指す。図5の例ではボイスチャット装置10-1~10-3が、P2P接続を行うボイスチャット装置10に相当する。また図4に示すボイスチャット装置10-1~10-5についても同様に、P2P接続を行うボイスチャット装置10に相当する。
そしてP2P接続を行わないボイスチャット装置10とは、パーティに参加する全部の他のユーザについて、当該ユーザが利用するボイスチャット装置10とは中継装置12を介して接続されるボイスチャット装置10を指す。図5の例ではボイスチャット装置10-4~10-6が、P2P接続を行わないボイスチャット装置10に相当する。また図3に示すボイスチャット装置10-1~10-5についても同様に、P2P接続を行わないボイスチャット装置10に相当する。
ここでは例えば、P2P接続を行うボイスチャット装置10を利用するユーザのユーザデータについては、P2P接続フラグの値が1に設定されることとする。また例えば、P2P接続を行わないボイスチャット装置10を利用するユーザのユーザデータについては、P2P接続フラグの値が0に設定されることとする。
図7には、6人のユーザが参加するパーティに対応付けられる、パーティIDが001であるパーティ管理データが例示されている。図7に示すパーティ管理データには、それぞれ当該パーティに参加するユーザに対応付けられる6個のユーザデータが含まれている。以下、ユーザIDがaaaであるユーザ、bbbであるユーザ、cccであるユーザ、dddであるユーザ、eeeであるユーザ、fffであるユーザを、それぞれ、ユーザA、ユーザB、ユーザC、ユーザD、ユーザE、ユーザFと呼ぶこととする。また、ユーザA、ユーザB、ユーザC、ユーザD、ユーザE、ユーザFは、それぞれ、ボイスチャット装置10-1、10-2、10-3、10-4、10-5、10-6を利用していることとする。
また本実施形態では、管理サーバ14に記憶されているパーティ管理データのコピーが、当該パーティ管理データに対応付けられるパーティに参加するユーザが利用するボイスチャット装置10、及び、中継装置12に送信される。そしてボイスチャット装置10の記憶部10b、及び、中継装置12の記憶部12bには、管理サーバ14に記憶されているパーティ管理データのコピーが記憶される。そのため、パーティに参加するユーザが利用するボイスチャット装置10は、当該パーティに参加する他のユーザが利用するボイスチャット装置10のアドレスを特定可能である。
また、ボイスチャット装置10の記憶部10bには、中継装置12のアドレスを示すデータも記憶されていることとする。そのため、ボイスチャット装置10は、中継装置12のアドレスを特定可能である。
また本実施形態では、例えばユーザによるパーティへの参加操作などに応じて、管理サーバ14に記憶されているパーティ管理データは更新される。そして管理サーバ14に記憶されているパーティ管理データが更新される度に、更新後のパーティ管理データのコピーが、当該パーティ管理データに対応付けられるパーティに参加するユーザが利用するボイスチャット装置10、及び、中継装置12に送信される。そして、ボイスチャット装置10の記憶部10b、及び、中継装置12の記憶部12bに記憶されているパーティ管理データのコピーは更新される。このようにして本実施形態では、パーティ管理データに示されている最新の情報が、当該パーティ管理データに対応付けられるパーティに参加するユーザが利用するボイスチャット装置10で共有されることとなる。
そして本実施形態では例えば、P2P接続を行うボイスチャット装置10は、P2P接続を行う他のボイスチャット装置10に宛てて送信する音声データを、当該ボイスチャット装置10に直接送信する。そして、P2P接続を行うボイスチャット装置10は、P2P接続を行わない他のボイスチャット装置10に宛てて送信する音声データを、中継装置12に送信する。
そして本実施形態では例えば、P2P接続を行わないボイスチャット装置10は、他のボイスチャット装置10に宛てて送信する音声データを、中継装置12に送信する。
このように図7に例示するパーティ管理データを用いることで、音声データをボイスチャット装置10に送信するか中継装置12に送信するかを適切に制御することができる。
また、以上で説明した音声データの選択において、当該音声データに対する音声区間検出処理の実行結果、又は、当該音声データが表す音声の音量の少なくとも一方に基づいて、複数の音声データのうちの一部が選択されるようにしてもよい。
例えば、ボイスチャット装置10は、所定の期間(例えば20ミリ秒、あるいは、40ミリ秒など)毎に、当該期間にわたって入力された音声をエンコードすることにより、当該期間に対応する音声データを生成するようにしてもよい。
そして当該ボイスチャット装置10が、当該音声データに対して、公知の音声区間検出(Voice Activity Detection(VAD))処理を実行することにより、当該音声データが人の声を表すものであるか否かを判定してもよい。そして、当該ボイスチャット装置10が、当該音声データが人の声を表すものであるか否かを示すVADデータを生成してもよい。ここで例えば、当該音声データが人の声を表すものである場合は、値が1であるVADデータが生成され、そうでない場合は、値が0であるVADデータが生成されてもよい。
また、当該ボイスチャット装置10が、当該音声データが表す音声の音量を特定してもよい。そして、当該ボイスチャット装置10が、当該音声データが表す音声の音量を示す音量データを生成してもよい。
そして当該ボイスチャット装置10は、当該ボイスチャット装置10の識別情報(例えば当該ボイスチャット装置10に対応するユーザID)、上述のVADデータ、及び、上述の音量データが関連付けられた、当該音声データを送信してもよい。また、タイムスタンプなどといった、当該音声データに対応する期間を示すデータが、当該音声データに関連付けられていてもよい。
そして、互いに異なるボイスチャット装置10からそれぞれ複数の音声データを受信する中継装置12は、所定の期間毎に、当該期間に受信した複数の音声データのうちの一部を選択するようにしてもよい。ここで例えば、音声データに関連付けられたVADデータや音量データに基づいて、当該選択が行われてもよい。そして、選択された一部の音声データが送信されるようにしてもよい。
また、互いに異なるボイスチャット装置10からそれぞれ複数の音声データを受信するボイスチャット装置10は、所定の期間毎に、当該期間に受信した複数の音声データのうちの一部を選択するようにしてもよい。ここで例えば、音声データに関連付けられたVADデータや音量データに基づいて、当該選択が行われてもよい。そして、選択された一部の音声データが当該ボイスチャット装置10のエンコード・デコード部10hに出力されるようにしてもよい。
VADデータや音量データに基づく選択の具体例については後述する。
以下、音声データの選択、及び、選択された音声データの送信を中心に、本実施形態に係るボイスチャットシステム1で実装される機能、及び、本実施形態に係るボイスチャットシステム1で実行される処理について、さらに説明する。
図8は、本実施形態に係るボイスチャット装置10及び中継装置12で実装される機能の一例を示す機能ブロック図である。なお、本実施形態に係るボイスチャット装置10及び中継装置12で、図8に示す機能のすべてが実装される必要はなく、また、図8に示す機能以外の機能が実装されていても構わない。
図8に示すように、本実施形態に係るボイスチャット装置10には、機能的には例えば、パーティ管理データ記憶部20、パーティ管理部22、音声受付部24、VADデータ生成部26、音量データ生成部28、音声データ送信部30、音声データ受信部32、選択部34、選択音声データ出力部36、音声出力部38、が含まれる。
パーティ管理データ記憶部20は、記憶部10bを主として実装される。パーティ管理部22は、プロセッサ10a及び通信部10cを主として実装される。音声受付部24は、マイク10f及びエンコード・デコード部10hを主として実装される。VADデータ生成部26、音量データ生成部28、選択部34、選択音声データ出力部36は、プロセッサ10aを主として実装される。音声データ送信部30、音声データ受信部32は、通信部10cを主として実装される。音声出力部38は、エンコード・デコード部10h及びスピーカ10gを主として実装される。
そして以上の機能は、コンピュータであるボイスチャット装置10にインストールされた、以上の機能に対応する指令を含むプログラムをプロセッサ10aで実行することにより実装されている。このプログラムは、例えば、光ディスク、磁気ディスク、磁気テープ、光磁気ディスク、フラッシュメモリ等のコンピュータ読み取り可能な情報記憶媒体を介して、あるいは、インターネットなどを介してボイスチャット装置10に供給される。
また、図8に示すように、本実施形態に係る中継装置12には、機能的には例えば、パーティ管理データ記憶部40、パーティ管理部42、音声データ受信部44、選択部46、音声データ送信部48、が含まれる。
パーティ管理データ記憶部40は、記憶部12bを主として実装される。パーティ管理部42は、プロセッサ12a及び通信部12cを主として実装される。音声データ受信部44、音声データ送信部48は、通信部12cを主として実装される。選択部46は、プロセッサ12aを主として実装される。
そして以上の機能は、コンピュータである中継装置12にインストールされた、以上の機能に対応する指令を含むプログラムをプロセッサ12aで実行することにより実装されている。このプログラムは、例えば、光ディスク、磁気ディスク、磁気テープ、光磁気ディスク、フラッシュメモリ等のコンピュータ読み取り可能な情報記憶媒体を介して、あるいは、インターネットなどを介して中継装置12に供給される。
ボイスチャット装置10のパーティ管理データ記憶部20、及び、中継装置12のパーティ管理データ記憶部40は、本実施形態では例えば、図7に例示するパーティ管理データを記憶する。
ボイスチャット装置10のパーティ管理部22は、本実施形態では例えば、管理サーバ14から送信されるパーティ管理データの受信に応じて、パーティ管理データ記憶部20に記憶されているパーティ管理データを受信したパーティ管理データに更新する。
中継装置12のパーティ管理部42は、本実施形態では例えば、管理サーバ14から送信されるパーティ管理データの受信に応じて、パーティ管理データ記憶部40に記憶されているパーティ管理データを受信したパーティ管理データに更新する。
例えば、ユーザが既存のパーティに対する参加操作を行うと、管理サーバ14は、当該パーティのパーティIDを含むパーティ管理データに、当該ユーザのユーザIDを含むユーザデータを追加する。以下、当該ユーザデータを追加ユーザデータと呼ぶこととする。ここで追加ユーザデータの接続先アドレスデータには、当該ユーザが利用するボイスチャット装置10のアドレスが設定される。また、追加ユーザデータの種類データには、当該ボイスチャット装置10の種類を示す値が設定される。
また追加ユーザデータのP2P接続フラグの値が、上述のように、1又は0に設定される。ここで例えば、ボイスチャットの音声を表す音声データが互いに送受信されるボイスチャット装置10の数に基づいて、当該ユーザデータのP2P接続フラグの値が決定されてもよい。例えば、当該パーティに対応するパーティ管理データに含まれるユーザデータの数に基づいて、当該ユーザデータのP2P接続フラグの値が決定されてもよい。
具体的には例えば、追加ユーザデータを含む、当該パーティ管理データに含まれているユーザデータの数が8以下である場合は、追加ユーザデータのP2P接続フラグの値が1に設定されてもよい。また、追加ユーザデータを含む、当該パーティ管理データに含まれているユーザデータの数が9以上である場合は、追加ユーザデータのP2P接続フラグの値が0に設定されてもよい。
このようにすれば、音声データが互いに送受信されるボイスチャット装置10が少ないうちはP2P接続により互いに音声データの送受信が行われる。この場合は、音声データの送受信に中継装置12は利用されない。そのため、音声データが互いに送受信されるボイスチャット装置10が少ないうちは中継装置12の負荷を抑えることができる。
一方で、音声データが互いに送受信されるボイスチャット装置10が多いと、1つのボイスチャット装置10の送信先となるボイスチャット装置10が多くなることから、コンピュータネットワーク16を流れる音声データの通信量は過度に多くなってしまう。ここで上述のように、音声データが互いに送受信されるボイスチャット装置10が多くなると中継装置12を利用した音声データの送受信が併せて行われるようにすることで、音声データの通信量が過度に多くなってしまうことを防ぐことができる。
なお、パーティ管理データに含まれているユーザデータの数が8以下である場合に、参加操作を行ったユーザが利用するボイスチャット装置10は、同じパーティの他のユーザが利用するボイスチャット装置10のそれぞれに対してP2P接続を試行してもよい。そしてすべてのボイスチャット装置10についてP2P接続が成功した場合に、追加ユーザデータのP2P接続フラグの値が1に設定されてもよい。一方で、いずれかのボイスチャット装置10についてP2P接続が失敗した場合に、追加ユーザデータのP2P接続フラグの値が0に設定されてもよい。
上述したように、このようにして管理サーバ14に記憶されているパーティ管理データに追加ユーザデータが追加されることに応じて、当該パーティに参加するユーザが利用するボイスチャット装置10に記憶されているパーティ管理データが更新される。また、中継装置12に記憶されているパーティ管理データも同様に更新される。
音声受付部24は、本実施形態では例えば、ボイスチャットの音声を受け付ける。音声受付部24は、当該音声をエンコードすることにより、当該音声を表す音声データを生成してもよい。
VADデータ生成部26は、本実施形態では例えば、音声受付部24が生成する音声データに基づいて、上述のVADデータを生成する。
音量データ生成部28は、本実施形態では例えば、音声受付部24が生成する音声データに基づいて、上述の音量データを生成する。
ボイスチャット装置10の音声データ送信部30は、本実施形態では例えば、音声受付部24が受け付ける音声を表す音声データを送信する。ここで音声データ送信部30は、当該ボイスチャット装置10の識別情報が関連付けられた音声データを送信してもよい。また、音声データ送信部30は、当該ボイスチャット装置10の識別情報、VADデータ生成部26により生成されるVADデータ、及び、音量データ生成部28により生成される音量データが関連付けられた音声データを送信してもよい。またタイムスタンプなどといった、当該音声データに対応する期間を示すデータが、当該音声データに関連付けられていてもよい。
また上述のように、ボイスチャット装置10の音声データ送信部30は、当該ボイスチャット装置10を利用するユーザと同じパーティに参加している他のユーザが利用するボイスチャット装置10に音声データを送信してもよい。また、ボイスチャット装置10の音声データ送信部30は、当該ボイスチャット装置10を利用するユーザと同じパーティに参加している他のユーザが利用するボイスチャット装置10に宛てた音声データを中継装置12に送信してもよい。
ボイスチャット装置10の音声データ受信部32は、本実施形態では例えば、音声データを受信する。ここで音声データ受信部32は、それぞれ互いに異なる送信装置から送信される複数の音声データを受信してもよい。上述の例では、当該ボイスチャット装置10を利用するユーザと同じパーティに参加している他のユーザが利用するボイスチャット装置10、又は、中継装置12が、当該送信装置に相当する。
ボイスチャット装置10の音声データ受信部32は、当該ボイスチャット装置10を利用するユーザと同じパーティに参加している他のユーザが利用するボイスチャット装置10から直接音声データを送信してもよい。また、ボイスチャット装置10の音声データ受信部32は、当該ボイスチャット装置10を利用するユーザと同じパーティに参加している他のユーザが利用するボイスチャット装置10から中継装置12を経由して送信される音声データを受信してもよい。
ボイスチャット装置10の選択部34は、本実施形態では例えば、ボイスチャット装置10の音声データ受信部32が受信した複数の音声データのうちの一部を選択する。ここで選択部34は、上述のように、音声データに対する音声区間検出処理の実行結果、又は、音声データが表す音声の音量の少なくとも一方に基づいて、複数の音声データのうちの一部を選択してもよい。
選択音声データ出力部36は、本実施形態では例えば、ボイスチャット装置10の選択部34により選択された一部の音声データを出力する。ここでは例えば、ボイスチャット装置10の選択部34により選択された一部の音声データは、音声出力部38に出力される。
音声出力部38は、本実施形態では例えば、選択音声データ出力部36から出力される音声データをデコードする。そして、音声出力部38は、本実施形態では例えば、当該音声データをデコードすることにより生成される、当該音声データを表す音声を出力する。
中継装置12の音声データ受信部44は、本実施形態では例えば、それぞれ互いに異なる送信装置から送信される複数の音声データを受信する。上述の例では、中継装置12に音声データを送信するボイスチャット装置10が、当該送信装置に相当する。中継装置12の音声データ受信部44は、例えば、ボイスチャット装置10の音声データ送信部30が送信する音声データを受信する。
中継装置12の選択部46は、本実施形態では例えば、中継装置12の音声データ受信部44が受信する複数の音声データのうちの一部を選択する。ここで選択部46は、上述のように、音声データに対する音声区間検出処理の実行結果、又は、音声データが表す音声の音量の少なくとも一方に基づいて、複数の音声データのうちの一部を選択してもよい。
中継装置12の音声データ送信部48は、本実施形態では例えば、中継装置12の選択部46により選択される一部の音声データを、中継装置12の音声データ受信部44が受信する音声データの送信元である送信装置とボイスチャットが可能な受信装置に送信する。上述の例では、ボイスチャット装置10が当該受信装置に相当する。ここで音声データ送信部48は、音声データに関連付けられているユーザIDに基づいて、当該ユーザIDが表すユーザが属するパーティを特定してもよい。そして、音声データ送信部48は、当該ユーザIDに対応付けられるユーザが利用するボイスチャット装置10以外の、当該パーティに参加するユーザが利用するボイスチャット装置10に、当該音声データを送信してもよい。
また、中継装置12は、それぞれ互いに異なる第1群に含まれる通信装置からそれぞれ送信される複数の音声データを受信してもよい。そして、中継装置12は、当該複数の音声データのうちから選択される一部を、当該音声データを送信した通信装置とは異なる通信装置に送信してもよい。図5及び図6の例では、ボイスチャット装置10-4~10-6が、第1群に含まれる通信装置に相当する。すなわち、P2P接続フラグの値が0であるユーザデータに対応付けられるボイスチャット装置10が第1群に含まれる通信装置に相当する。
また、第2群に含まれるボイスチャット装置10が、中継装置12から送信される上述の一部の音声データ、及び、それぞれ互いに異なる第2群に含まれる他の通信装置から送信される少なくとも1つの音声データを受信してもよい。そして、当該ボイスチャット装置10は、受信する複数の音声データのうちの一部を、当該ボイスチャット装置10の音声出力部38に出力してもよい。図6の例では、ボイスチャット装置10-1が、音声データを受信するボイスチャット装置10に相当する。また、ボイスチャット装置10-1~10-3が、第2群に含まれる通信装置に相当する。すなわち、P2P接続フラグの値が1であるユーザデータに対応付けられるボイスチャット装置10が第2群に含まれる通信装置に相当する。
また上述のように、管理サーバ14が、音声データが互いに送受信される通信装置の数に基づいて、ボイスチャットシステム1に新たに追加される通信装置を第1群に含めるか第2群に含めるかを決定してもよい。例えば上述のように、管理サーバ14が、音声データが互いに送受信されるボイスチャット装置10の数に基づいて、新たに追加されるボイスチャット装置10に対応するユーザデータに含まれるP2P接続フラグの値を決定してもよい。
ここで、本実施形態に係るボイスチャット装置10において行われる、音声データの送信処理の流れの一例を、図9に例示するフロー図を参照しながら説明する。図9に示すS101~S107に示す処理は、所定の期間毎(例えば20ミリ秒毎、あるいは、40ミリ秒毎)に繰り返し実行される。
まず、音声受付部24が、本ループの期間において受け付けた音声をエンコードすることにより、音声データを生成する(S101)。
そして、VADデータ生成部26が、S101に示す処理で生成された音声データに対して、VAD処理を実行することにより、当該音声データが人の声を表すものであるか否かを判定する(S102)。
そして、VADデータ生成部26が、S102に示す処理での判定結果に応じたVADデータを生成する(S103)。
そして、音量データ生成部28が、S101に示す処理で生成された音声データが表す音声の音量を特定する(S104)。
そして、音量データ生成部28が、S104に示す処理で特定された音量を示す音量データを生成する(S105)。
そして、音声データ送信部30が、パーティ管理データ記憶部20に記憶されているパーティ管理データに基づいて、音声データの送信先である通信装置を特定する(S106)。ここでは例えば、送信先となるボイスチャット装置10のアドレスや、中継装置12への音声データの送信要否などが特定される。
そして、音声データ送信部30が、S106に示す処理で特定された送信先にS101に示す処理で生成された音声データを送信して(S107)、S101に示す処理に戻る。上述のように、当該音声データには、当該ボイスチャット装置10を利用するユーザのユーザID、S103に示す処理で生成されたVADデータ、及び、S105に示す処理で生成された音量データが関連付けられている。またタイムスタンプなどといった、本ループの期間を示すデータが、当該音声データに関連付けられていてもよい。
なお、チャタリングを防止するために、S102に示す処理で音声データが人の声を表すものであると判定された際には、判定されたタイミングから所定の時間(例えば1秒)にわたって、S102に示す処理が実行されないようにしてもよい。そして当該時間にわたって、S103に示す処理では値が1であるVADデータが生成されてもよい。
次に、本実施形態に係るボイスチャット装置10において行われる、音声の出力処理の流れの一例を、図10A及び図10Bに例示するフロー図を参照しながら説明する。図10A及び図10Bに示すS201~S217に示す処理は、所定の期間毎(例えば20ミリ秒毎、あるいは、40ミリ秒毎)に繰り返し実行される。また、以下の説明では、1つの期間において当該ボイスチャット装置10において音声出力部38が入力可能な音声データの最大数は予め定められていることとする。以下、当該最大数をn1と表現する。
まず、選択部34が、本ループの期間において音声データ受信部32が受信した音声データを特定する(S201)。
そして、選択部34が、S201に示す処理で特定された音声データの数(以下、m1と表現する)がn1以下であるか否かを確認する(S202)。
S202に示す処理においてm1がn1以下であることが確認されたとする(S202:Y)。この場合は、選択音声データ出力部36は、S201に示す処理で特定されたすべての音声データを音声出力部38に出力する(S203)。
S202に示す処理においてm1がn1より大きいことが確認されたとする(S202:N)。この場合は、選択部34は、S201に示す処理で特定された複数の音声データのうちから、関連付けられているVADデータの値が1である音声データを特定する(S204)。
そして、選択部34は、S204に示す処理で特定された音声データの数(以下、m2と表現する)がn1以上であるか否かを確認する(S205)。
S205に示す処理においてm2がn1以上であることが確認されたとする(S205:Y)。この場合は、選択部34は、S204に示す処理で特定された複数の音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn1個の音声データを特定する(S206)。m2がn1と同じである場合は、S206に示す処理において、S204に示す処理で特定されたすべての音声データが特定されることとなる。
そして、選択音声データ出力部36は、S206に示す処理で特定されたn1個の音声データを音声出力部38に出力する(S207)。
S205に示す処理においてm2がn1より小さいことが確認されたとする(S205:N)。この場合は、選択部34は、S201に示す処理で特定された複数の音声データのうちから、関連付けられているVADデータの値が0である音声データを特定する(S208)。
そして、選択部34は、S208に示す処理で特定された音声データのそれぞれについて、当該音声データが表す音声の音量の移動平均を特定する(S209)。ここで音声データが表す音声の音量の移動平均とは、例えば、当該音声データの送信元であるボイスチャット装置10から、直近の所定回あるいは直近の所定時間(例えば直近の1秒間)に受信した音声データが表す音声の音量の平均を指す。なお移動平均を特定するにあたり、選択部34が、直近の所定回あるいは直近の所定時間にわたって受信した音声データを少なくとも記憶するようにしてもよい。
そして、選択部34は、S208に示す処理で特定された複数の音声データのうちから、S209に示す処理で特定された移動平均が所定の閾値以上(例えば-40dBOV以上)である音声データを特定する(S210)。
そして、選択部34は、S204に示す処理で特定された音声データの数と、S210に示す処理で特定された音声データの数との合計(以下、m3と表現する)と、n1とを比較する(S211)。
S211に示す処理においてm3がn1より大きいことが確認されたとする。この場合は、選択部34は、S204又はS210に示す処理で特定された複数の音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn1個の音声データを特定する(S212)。そして、選択音声データ出力部36は、S212に示す処理で特定されたn1個の音声データを音声出力部38に出力する(S213)。
S211に示す処理においてm3がn1より小さいことが確認されたとする。この場合は、選択部34は、S204又はS210に示す処理で特定されていない残りの音声データのうちから、関連付けられている音量データが示す音量が大きいものから順に(n1-m3)個の音声データを特定する(S214)。そして、選択音声データ出力部36は、S204、S210、又は、S214に示す処理で特定された合計n1個の音声データを音声出力部38に出力する(S215)。
S211に示す処理においてm3がn1と同じであることが確認されたとする。この場合は、選択部34は、S204又はS210に示す処理で特定された合計n1個の音声データを音声出力部38に出力する(S216)。
そして音声出力部38は、S203、S207、S213、S215、又は、S216に示す処理で出力された音声データをデコードして、当該音声データが表す音声を出力する(S217)。そして、S201に示す処理に戻る。
図10A及び図10Bに示す処理例によれば、ボイスチャット装置10の選択部34によって、複数の音声データのうちから、ボイスチャットの音声を表すものである可能性が高い音声データが的確に選択されることとなる。
なお、音声データの選択処理は、図10A及び図10Bに示す処理例に示されている処理には限定されない。例えば、選択部34は、関連付けられているVADデータの値が1である、あるいは、音量の移動平均が所定の閾値以上である、音声データを特定してもよい。そして、選択部34は、これらの音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn1個の音声データを特定してもよい。
次に、本実施形態に係る中継装置12において行われる、音声データの中継処理の流れの一例を、図11A及び図11Bに例示するフロー図を参照しながら説明する。図11A及び図11Bに示すS301~S316に示す処理は、所定の期間毎(例えば20ミリ秒毎、あるいは、40ミリ秒毎)に繰り返し実行される。また、以下の説明では、1つの期間において中継装置12において送信可能な音声データの最大数は予め定められていることとする。以下、当該最大数をn2と表現する。
また中継装置12の選択部46は、予め、パーティ管理データ記憶部40に記憶されているパーティ管理データに基づいて、音声データの送信先である複数のボイスチャット装置10を特定していることとする。
まず、選択部46が、本ループの期間において音声データ受信部44が受信した音声データを特定する(S301)。
そして、選択部46が、S301に示す処理で特定された音声データの数(以下、m4と表現する)がn2以下であるか否かを確認する(S302)。
S302に示す処理においてm4がn2以下であることが確認されたとする(S302:Y)。この場合は、音声データ送信部48は、S301に示す処理で特定されたすべての音声データを、送信先であるボイスチャット装置10に送信して(S303)、S301に示す処理に戻る。
S302に示す処理においてm4がn2より大きいことが確認されたとする(S302:N)。この場合は、選択部46は、S301に示す処理で特定された複数の音声データのうちから、関連付けられているVADデータの値が1である音声データを特定する(S304)。
そして、選択部46は、S304に示す処理で特定された音声データの数(以下、m5と表現する)がn2以上であるか否かを確認する(S305)。
S305に示す処理においてm5がn2以上であることが確認されたとする(S305:Y)。この場合は、選択部46は、S204に示す処理で特定された複数の音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn2個の音声データを特定する(S306)。m5がn2と同じである場合は、S306に示す処理において、S304に示す処理で特定されたすべての音声データが特定されることとなる。
そして、音声データ送信部48は、S306に示す処理で特定されたn2個の音声データを、送信先であるボイスチャット装置10に送信して(S307)、S301に示す処理に戻る。
S305に示す処理においてm5がn2より小さいことが確認されたとする(S305:N)。この場合は、選択部46は、S301に示す処理で特定された複数の音声データのうちから、関連付けられているVADデータの値が0である音声データを特定する(S308)。
そして、選択部46は、S308に示す処理で特定された音声データのそれぞれについて、当該音声データが表す音声の音量の移動平均を特定する(S309)。ここで音声データが表す音声の音量の移動平均とは、例えば、当該音声データの送信元であるボイスチャット装置10から、直近の所定回あるいは直近の所定時間(例えば直近の1秒間)に受信した音声データが表す音声の音量の平均を指す。なお移動平均を特定するにあたり、選択部46が、直近の所定回あるいは直近の所定時間にわたって受信した音声データを少なくとも記憶するようにしてもよい。
そして、選択部46は、S308に示す処理で特定された複数の音声データのうちから、S309に示す処理で特定された移動平均が所定の閾値以上(例えば-40dBOV以上)である音声データを特定する(S310)。
そして、選択部46は、S304に示す処理で特定された音声データの数と、S210に示す処理で特定された音声データの数との合計(以下、m6と表現する)と、n2とを比較する(S311)。
S311に示す処理においてm6がn2より大きいことが確認されたとする。この場合は、選択部46は、S304又はS310に示す処理で特定された複数の音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn2個の音声データを特定する(S312)。そして、音声データ送信部48は、S312に示す処理で特定されたn2個の音声データを、送信先であるボイスチャット装置10に送信して(S313)、S301に示す処理に戻る。
S311に示す処理においてm6がn2より小さいことが確認されたとする。この場合は、選択部46は、S304又はS310に示す処理で特定されていない残りの音声データのうちから、関連付けられている音量データが示す音量が大きいものから順に(n2-m6)個の音声データを特定する(S314)。そして、選択音声データ出力部36は、S304、S310、又は、S314に示す処理で特定された合計n2個の音声データを、送信先であるボイスチャット装置10に送信して(S315)、S301に示す処理に戻る。
S311に示す処理においてm6がn2と同じであることが確認されたとする。この場合は、選択部46は、S304又はS310に示す処理で特定された合計n2個の音声データを、送信先であるボイスチャット装置10に送信して(S316)、S301に示す処理に戻る。
図11A及び図11Bに示す処理例によれば、中継装置12の選択部46によって、複数の音声データのうちから、ボイスチャットの音声を表すものである可能性が高い音声データが的確に選択されることとなる。
VADデータに基づく音声データの選択を行うことで、例えば、机をたたく音や救急車の音などといった人の声以外の音を表す音声データが選択される可能性を下げることができる。また、音量の移動平均に基づく音声データの選択を行うことで、例えば、VADデータに基づく音声データの選択では選択されなかったが実際には人の声を表す音声データが選択される可能性を上げることができる。
なお、音声データの選択処理は、図11A及び図11Bに示す処理例に示されている処理には限定されない。例えば、選択部46は、関連付けられているVADデータの値が1である、あるいは、音量の移動平均が所定の閾値以上である、音声データを特定してもよい。そして、選択部46は、これらの音声データのうちから、関連付けられている音量データが示す音量が大きいものから順にn2個の音声データを特定してもよい。
また上述のように、選択部34、又は、選択部46は、音声データが表す音声の音量の移動平均に基づいて、複数の音声データのうちの一部を選択してもよい。このようにすれば、音声データの選択が安定して行われることとなる。
また図11A及び図11Bに示す処理例において、選択部46は、複数の音声データのうちから、音声データの送信先であるボイスチャット装置10の種類に基づいて決定される数の音声データを選択してもよい。そして、音声データ送信部48は、当該一部の音声データを当該ボイスチャット装置10に送信するようにしてもよい。この場合は、上述の数n2は、送信先であるボイスチャット装置10の種類によって異なることとなる。
スマートフォンは、キャリアのネットワークを利用する可能性が高いため、当該ネットワークを流れる音声データの通信量は特に抑える必要がある。このことを踏まえると、例えば、スマートフォンであるボイスチャット装置10には、ゲームコンソールであるボイスチャット装置10よりも少ない数の音声データが送信されるようにすることが好適である。
また、複数の音声データのうちから、装置の負荷又はコンピュータネットワーク16の通信品質に基づいて決定される数の音声データが選択されるようにしてもよい。
例えば、ボイスチャット装置10の選択部34は、当該ボイスチャット装置10の負荷に基づいて、上述の値n1を決定してもよい。
また、中継装置12の選択部46は、中継装置12の負荷、又は、コンピュータネットワーク16の通信品質(例えば通信量)に基づいて、上述の値n2を決定してもよい。
以上で説明したボイスチャットシステム1によれば、ボイスチャットの音声を表す音声データの出力を適切に間引くことができる。例えば、中継装置12による音声データの送信出力や、ボイスチャット装置10のプロセッサ10aからエンコード・デコード部10hへの音声データの出力を、適切に間引くことができる。
なお、本発明は上述の実施形態に限定されるものではない。
例えば図10A及び図10Bに示す処理において、選択部34が、最近発話したユーザを表すユーザIDのリストを記憶してもよい。当該リストには、n1個のユーザIDが含まれていてもよい。
そして、選択部34が、リストにユーザIDが含まれていないユーザが利用するボイスチャット装置10が送信する音声データのうちの一部を選択してもよい。ここでは選択される音声データの数n3は、n1よりも小さい。そして、選択部34は、リストからn3個のユーザIDを削除してもよい。例えば、リストに追加されたタイミングが最も古いものから順にn3個のユーザIDがリストから削除されるようにしてもよい。また、直近における音声データが表す音声の音量が小さいものから順にn3個のユーザIDがリストから削除されるようにしてもよい。
そして、選択部34は、選択されたn3個の音声データに対応付けられるユーザIDを当該リストに追加してもよい。
そして選択部34は、リストに含まれるユーザIDが関連付けられている音声データを選択してもよい。このようにすれば、ボイスチャット装置10における音声データの選択が安定して行われることとなる。
また、同様に、図11A及び図11Bに示す処理において、選択部46が、n2個のユーザIDを含むリストを記憶してもよい。そして上述の例における選択部34と同様にして、当該リストに含まれるユーザIDが関連付けられている音声データが選択されるようにしてもよい。このようにすれば、中継装置12における音声データの選択が安定して行われることとなる。
また、ボイスチャット装置10、及び、中継装置12の役割分担は上述のものに限定されない。例えば、中継装置12において、音声データが人の声を表すものであるか否かの判定や、音声データが表す音声の音量の特定が実行されてもよい。また、例えば、音声データを受信するボイスチャット装置10において、音声データが人の声を表すものであるか否かの判定や、音声データが表す音声の音量の特定が実行されてもよい。
また、上記の具体的な文字列や数値及び図面中の具体的な文字列や数値は例示であり、これらの文字列や数値には限定されない。
Claims (9)
- それぞれ互いに異なる送信装置から送信される複数の音声データを受信する受信部と、
前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択する選択部と、
選択される前記一部の前記音声データを出力する出力部と、
を含むことを特徴とする音声出力制御装置。 - 前記出力部は、選択される前記一部の前記音声データを、前記送信装置とボイスチャットが可能な受信装置に送信する、
ことを特徴とする請求項1に記載の音声出力制御装置。 - 前記選択部は、前記複数の音声データのうちから、前記受信装置の種類に基づいて決定される数の前記音声データを選択する、
ことを特徴とする請求項2に記載の音声出力制御装置。 - 前記音声データをデコードするデコード部、をさらに含み、
前記出力部は、選択される前記一部の前記音声データを前記デコード部に出力する、
ことを特徴とする請求項1に記載の音声出力制御装置。 - 前記選択部は、前記複数の音声データのうちから、前記音声出力制御装置の負荷又は前記音声出力制御装置が接続されているコンピュータネットワークの通信品質に基づいて決定される数の前記音声データを選択する、
ことを特徴とする請求項1から4のいずれか一項に記載の音声出力制御装置。 - 第1群に含まれる複数の通信装置と、第2群に含まれる複数の通信装置と、中継装置と、を含み、
前記中継装置は、
それぞれ互いに異なる前記第1群に含まれる前記通信装置から送信される複数の音声データを受信する第1受信部と、
前記中継装置の前記第1受信部が受信する前記複数の前記音声データのうちから選択される一部を、当該音声データを送信した前記通信装置とは異なる前記通信装置に送信する送信部と、を含み、
前記第2群に含まれる前記通信装置は、
前記音声データをデコードするデコード部と、
前記中継装置から送信される前記一部の前記音声データ、及び、それぞれ互いに異なる前記第2群に含まれる他の前記通信装置から送信される少なくとも1つの音声データを受信する第2受信部と、
前記通信装置の前記第2受信部が受信する複数の前記音声データのうちの一部を前記デコード部に出力する出力部と、
を含むことを特徴とする音声出力制御システム。 - 前記音声データが互いに送受信される前記通信装置の数に基づいて、前記音声出力制御システムに新たに追加される前記通信装置を前記第1群に含めるか前記第2群に含めるかを決定する決定部、をさらに含む、
ことを特徴とする請求項6に記載の音声出力制御システム。 - それぞれ互いに異なる送信装置から送信される複数の音声データを受信するステップと、
前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択するステップと、
選択される前記一部の前記音声データを出力するステップと、
を含むことを特徴とする音声出力制御方法。 - それぞれ互いに異なる送信装置から送信される複数の音声データを受信する手順、
前記音声データに対する音声区間検出処理の実行結果、又は、前記音声データが表す音声の音量の移動平均の少なくとも一方に基づいて、前記複数の音声データのうちの一部を選択する手順、
選択される前記一部の前記音声データを出力する手順、
をコンピュータに実行させることを特徴とするプログラム。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| US17/429,656 US12033655B2 (en) | 2019-02-19 | 2020-02-13 | Sound output control apparatus, sound output control system, sound output control method, and program |
| JP2021501923A JP7116240B2 (ja) | 2019-02-19 | 2020-02-13 | 音声出力制御システム、中継装置、通信装置、音声出力制御方法及びプログラム |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2019027628 | 2019-02-19 | ||
| JP2019-027628 | 2019-02-19 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2020170946A1 true WO2020170946A1 (ja) | 2020-08-27 |
Family
ID=72144848
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2020/005634 Ceased WO2020170946A1 (ja) | 2019-02-19 | 2020-02-13 | 音声出力制御装置、音声出力制御システム、音声出力制御方法及びプログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US12033655B2 (ja) |
| JP (1) | JP7116240B2 (ja) |
| WO (1) | WO2020170946A1 (ja) |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101252452A (zh) * | 2007-03-31 | 2008-08-27 | 红杉树(杭州)信息技术有限公司 | 一种多媒体会议中分布式混音系统 |
| EP2018058A1 (en) * | 2007-07-17 | 2009-01-21 | Huawei Technologies Co., Ltd. | Method for displaying speaker in video conference and device and system thereof |
| CN104167210A (zh) * | 2014-08-21 | 2014-11-26 | 华侨大学 | 一种轻量级的多方会议混音方法和装置 |
| US20150281648A1 (en) * | 2014-03-31 | 2015-10-01 | Polycom, Inc. | System and method for a hybrid topology media conferencing system |
| US20180091564A1 (en) * | 2016-09-28 | 2018-03-29 | Atlassian Pty Ltd | Dynamic adaptation to increased sfu load by disabling video streams |
Family Cites Families (13)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5978463A (en) * | 1997-04-18 | 1999-11-02 | Mci Worldcom, Inc. | Reservation scheduling system for audio conferencing resources |
| US6798786B1 (en) * | 1999-06-07 | 2004-09-28 | Nortel Networks Limited | Managing calls over a data network |
| US6665728B1 (en) * | 1998-12-30 | 2003-12-16 | Intel Corporation | Establishing optimal latency in streaming data applications that use data packets |
| US6839416B1 (en) * | 2000-08-21 | 2005-01-04 | Cisco Technology, Inc. | Apparatus and method for controlling an audio conference |
| US7221663B2 (en) * | 2001-12-31 | 2007-05-22 | Polycom, Inc. | Method and apparatus for wideband conferencing |
| US7773581B2 (en) * | 2004-03-19 | 2010-08-10 | Ericsson Ab | Method and apparatus for conferencing with bandwidth control |
| US20110044474A1 (en) * | 2009-08-19 | 2011-02-24 | Avaya Inc. | System and Method for Adjusting an Audio Signal Volume Level Based on Whom is Speaking |
| US20140369528A1 (en) * | 2012-01-11 | 2014-12-18 | Google Inc. | Mixing decision controlling decode decision |
| US9197848B2 (en) * | 2012-06-25 | 2015-11-24 | Intel Corporation | Video conferencing transitions among a plurality of devices |
| US8948058B2 (en) * | 2012-07-23 | 2015-02-03 | Cisco Technology, Inc. | System and method for improving audio quality during web conferences over low-speed network connections |
| CN104486518B (zh) | 2014-12-03 | 2017-06-30 | 中国电子科技集团公司第三十研究所 | 一种带宽受限网络环境下的电话会议分布式混音方法 |
| CN105721469B (zh) * | 2016-02-18 | 2019-09-20 | 腾讯科技(深圳)有限公司 | 音频数据处理方法、服务器、客户端以及系统 |
| US10375131B2 (en) * | 2017-05-19 | 2019-08-06 | Cisco Technology, Inc. | Selectively transforming audio streams based on audio energy estimate |
-
2020
- 2020-02-13 WO PCT/JP2020/005634 patent/WO2020170946A1/ja not_active Ceased
- 2020-02-13 JP JP2021501923A patent/JP7116240B2/ja active Active
- 2020-02-13 US US17/429,656 patent/US12033655B2/en active Active
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN101252452A (zh) * | 2007-03-31 | 2008-08-27 | 红杉树(杭州)信息技术有限公司 | 一种多媒体会议中分布式混音系统 |
| EP2018058A1 (en) * | 2007-07-17 | 2009-01-21 | Huawei Technologies Co., Ltd. | Method for displaying speaker in video conference and device and system thereof |
| US20150281648A1 (en) * | 2014-03-31 | 2015-10-01 | Polycom, Inc. | System and method for a hybrid topology media conferencing system |
| CN104167210A (zh) * | 2014-08-21 | 2014-11-26 | 华侨大学 | 一种轻量级的多方会议混音方法和装置 |
| US20180091564A1 (en) * | 2016-09-28 | 2018-03-29 | Atlassian Pty Ltd | Dynamic adaptation to increased sfu load by disabling video streams |
Also Published As
| Publication number | Publication date |
|---|---|
| US12033655B2 (en) | 2024-07-09 |
| JPWO2020170946A1 (ja) | 2021-11-18 |
| JP7116240B2 (ja) | 2022-08-09 |
| US20220208210A1 (en) | 2022-06-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10155164B2 (en) | Immersive audio communication | |
| RU2495538C2 (ru) | Устройства и способы для использования в создании аудиосцены | |
| US7090582B2 (en) | Use of multiple player real-time voice communications on a gaming device | |
| JP2008547290A5 (ja) | ||
| JP5122645B2 (ja) | 送信者に選択されている音声項目の対話参加者への提供方法 | |
| WO2023024894A1 (zh) | 一种多设备同步播放方法及装置 | |
| JP7143874B2 (ja) | 情報処理装置、情報処理方法およびプログラム | |
| CN115485705A (zh) | 信息处理装置、信息处理方法及程序 | |
| CN113300934B (zh) | 通信方法、装置、设备和存储介质 | |
| JP7232846B2 (ja) | ボイスチャット装置、ボイスチャット方法及びプログラム | |
| JP2026026128A (ja) | 音声送受信システム | |
| JP7116240B2 (ja) | 音声出力制御システム、中継装置、通信装置、音声出力制御方法及びプログラム | |
| JP5602688B2 (ja) | 音像定位制御システム、コミュニケーション用サーバ、多地点接続装置、及び音像定位制御方法 | |
| JP6473203B1 (ja) | サーバ装置、制御方法及びプログラム | |
| JP2024123488A (ja) | プログラム、情報処理装置の制御方法、及び情報処理装置 | |
| CN115309304A (zh) | 一种会话消息显示方法、装置、存储介质及计算机设备 | |
| WO2024190059A1 (ja) | システム、方法、およびプログラム | |
| KR100587772B1 (ko) | 음성 통신 방법 및 시스템 | |
| JP2025001682A (ja) | 遠隔会議システム及び遠隔会議方法 | |
| JP2022179354A (ja) | 情報処理装置およびプログラム | |
| AU2012202422B2 (en) | Immersive Audio Communication | |
| KR20250155763A (ko) | 화상회의 참여자간 음악 공유 방법 및 장치 | |
| KR101637208B1 (ko) | 모션이모티콘 제공 시스템, 방법 및 컴퓨터 프로그램이 기록된 기록매체 | |
| AU2006261594A1 (en) | Immersive audio communication |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 20759787 Country of ref document: EP Kind code of ref document: A1 |
|
| ENP | Entry into the national phase |
Ref document number: 2021501923 Country of ref document: JP Kind code of ref document: A |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 20759787 Country of ref document: EP Kind code of ref document: A1 |