WO2014115222A1 - Sound signal description method, sound signal production equipment, and sound signal reproduction equipment - Google Patents

Sound signal description method, sound signal production equipment, and sound signal reproduction equipment Download PDF

Info

Publication number
WO2014115222A1
WO2014115222A1 PCT/JP2013/007390 JP2013007390W WO2014115222A1 WO 2014115222 A1 WO2014115222 A1 WO 2014115222A1 JP 2013007390 W JP2013007390 W JP 2013007390W WO 2014115222 A1 WO2014115222 A1 WO 2014115222A1
Authority
WO
WIPO (PCT)
Prior art keywords
sound
sound field
sound signal
layered
reproduction
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2013/007390
Other languages
French (fr)
Inventor
Kaoru Watanabe
Satoshi OODE
Ikuko SAWAYA
Jae-Hyoun Yoo
Taejin Lee
Kyeongok Kang
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Japan Broadcasting Corp
Electronics and Telecommunications Research Institute ETRI
Original Assignee
Nippon Hoso Kyokai NHK
Japan Broadcasting Corp
Electronics and Telecommunications Research Institute ETRI
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nippon Hoso Kyokai NHK, Japan Broadcasting Corp, Electronics and Telecommunications Research Institute ETRI filed Critical Nippon Hoso Kyokai NHK
Priority to US14/652,907 priority Critical patent/US20150334502A1/en
Priority to KR1020157018270A priority patent/KR101682323B1/en
Publication of WO2014115222A1 publication Critical patent/WO2014115222A1/en
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S3/00Systems employing more than two channels, e.g. quadraphonic

Definitions

  • the present invention relates to a sound signal description method, a sound signal production equipment, and a sound signal reproduction equipment, all of which are capable of representing information of sound signals with use of metadata for sound reproduction through multichannel speakers.
  • ITU-R which is an international standardization body associated with broadcasting including sound, has defined requirements for an advanced multichannel sound system as ITU-R Recommendation. (Refer to Non Patent Literature 1.)
  • NPL 1 Performance requirements for an advanced multichannel stereophonic sound system for use with or without accompanying picture, Recommendation ITU-R BS.1909.
  • the format of "sound signals to compose a multi-layered sound field" can be used so as to facilitate rendering, conversion, and switching of received sound signals according to a receiver's environment or demand of program exchange or a home reproduction.
  • the receiver of program exchange or the home sometimes does not employ the same size image display as in the program production, and according to such a video reproduction environment of the receiver, the sound signal needs to be converted.
  • the study has not been conducted on the description method for the "sound signals to compose a multi-layered sound field.”
  • the present invention has been conceived in view of the above problems and aims to provide a sound signal description method corresponding to the format of the "sound signals to compose a multi-layered sound field", as well as a sound signal production equipment and a sound signal reproduction equipment which correspond to the sound signal description method.
  • one aspect of the present invention provides a sound signal description method for describing a multi-layered sound field, comprising: the number of sound field layers of the multi-layered sound field; a type of each sound field layer of the multi-layered sound field; and language information.
  • each sound field layer of the multi-layered sound field indicates the sound elements of the program, such as one of international sound, which consists of all the sound program elements except for the commentary/dialogue elements, and one of commentary/dialogue sound with particular language.
  • another aspect of the present invention provides a sound signal description method for describing a multi-layered sound field, comprising: the number of sound field layers of the multi-layered sound field; and a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
  • yet another aspect of the present invention provides a sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information; a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and a multiplexer that multiplexes the produced sound signal into a bit stream.
  • yet another aspect of the present invention provides a sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: an environment information input unit that inputs reproduction environment and user demand information; and a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information included in the sound signal and according to the reproduction environment and user demand information, and reproduces the converted sound signal.
  • each sound field layer of the multi-layered sound field indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language, and the particular language being switched by the environment information input unit.
  • the rendering reproduction unit preferably adds the sound signal of the particular language to the international sound and reproduces added sound.
  • yet another aspect of the present invention provides a sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field and a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video; a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and a multiplexer that multiplexes the produced sound signal into a bit stream.
  • yet another aspect of the present invention provides a sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: an environment information input unit that inputs reproduction environment information into the reproduction equipment; and a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field and a video link identifier included in the sound signal and according to the reproduction environment information.
  • the video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
  • the rendering reproduction unit preferably renders the sound signal of the sound field layer based on video display information input by the environment information input unit.
  • the sound signal description method, the sound signal production equipment, and the sound signal reproduction equipment according to the present invention make it possible to describe the "sound signals to compose a multi-layered sound field" and to produce and reproduce a sound program using the sound signals.
  • FIG. 1 shows an exemplary structure of an "Extended sound field descriptor" according to one embodiment of the present invention.
  • FIG. 2 shows a block diagram of a sound signal production equipment according to one embodiment of the present invention.
  • FIG. 3 shows a block diagram of a sound signal reproduction equipment according to one embodiment of the present invention.
  • FIG. 4 is a conceptual diagram of a multi-layered sound field in connection with narration language switching.
  • FIG. 5 shows a difference in display size between a program production environment and a reproduction environment.
  • FIG. 6 is a conceptual diagram of the multi-layered sound field associated with linked/unlinked video and sound.
  • FIG. 7 shows an exemplary structure of a "Basic sound field descriptor".
  • the present invention extends a description method (referred to below as a "Basic sound field descriptor") for describing "sound signals to compose a single-layered sound field” to the description method (referred to below as an "Extended sound field descriptor") for describing a "sound signals to compose a multi-layered sound field.”
  • a description method referred to below as a "Basic sound field descriptor”
  • extended sound field descriptor for describing a "sound signals to compose a multi-layered sound field.
  • descriptor which is described as metadata in a header of a corresponding multichannel sound signal or in the headers on each sound channel constituting the multichannel.
  • Table 1 illustrates terms and definitions of the Basic sound field descriptor.
  • the Basic sound field descriptor is employed for production and exchange of complete mix programs (i.e. programs including all sound required for reproduction) with multichannel sound, for example.
  • the Sound Essence descriptor includes a descriptor of a program, a descriptor (name) of the Sound-field, and other relevant descriptors.
  • the Sound-field is described by the Sound-field configuration with a hierarchical structure.
  • the Sound Channel descriptor includes the Channel label descriptor and/or Channel Position descriptor.
  • the Basic sound field descriptor includes (A) Sound Essence descriptors, (B) Sound-field configuration descriptors, and (C) Sound Channel descriptors.
  • Table 2 shows (A) Sound Essence descriptors in the Basic sound field descriptor.
  • Table 3 shows (B) Sound-field configuration descriptors in the Basic sound field descriptor.
  • Table 4 shows (C) Sound Channel descriptors in the Basic sound field descriptor.
  • Table 5 shows C.1 Channel label descriptors, which are descriptors of the Channel label data included in the Sound Channel descriptors.
  • Table 6 shows C.2 Channel position descriptors, which are descriptors of the Channel position data included in the Sound Channel descriptors.
  • the present invention extends the Basic sound field descriptor, which is the description method for the "sound signals to compose a single-layered sound field" as mentioned above, to the Extended sound field descriptor, which is the description method for the "sound signals to compose a multi-layered sound field.”
  • Table 7 illustrates terms and definitions of the Extended sound field descriptor.
  • the Sound Essence descriptor includes the descriptor of the program, the descriptor (name) of the Sound-field, and the other relevant descriptors.
  • the Sound-field in the Extended sound field descriptor is described by multiple Sound-field configurations (Group of sound-field configurations) (Sound space configurations) each having the hierarchical structure.
  • the Sound Channel descriptor includes the Channel label descriptor and/or the Channel Position descriptor.
  • Table 8 shows (A) Sound Essence descriptors in the Extended sound field descriptor.
  • Table 9 shows A.2 Sound-field descriptors in the Extended sound field descriptor.
  • FIG. 2 shows a block diagram of a sound signal production equipment according to one embodiment of the present invention.
  • the sound signal production equipment In order to "facilitate" rendering, conversion, and switching of received sound signals according to the receiver's environment or demand of program exchange or the home reproduction, the sound signal production equipment produces a sound program according to the Extended sound field descriptor, which is the format of the "sound signals to compose a multi-layered sound field.”
  • the sound signal production equipment inserts the Extended sound field descriptor as metadata into the header of the corresponding sound format signal or into the header of each audio signal, for program exchange and transmission to the home.
  • the sound signal production equipment includes a mixing unit 11, a metadata addition unit 12, a coding unit 13, a multiplexer 14, and a monitoring unit 15.
  • the mixing unit 11 mixes sound signals (Sound Sources 1-M) and outputs, to the coding unit 13, sound signals to compose the multi-layered sound field including Spatial anchor, Commentary, Dialogue, and Object signals, the sound signals being output from a "production system for sound signals to compose a multi-layered sound field.”
  • the metadata addition unit 12 outputs, to the coding unit 13, the metadata to be described for the Extended sound field descriptor of the multi-layered sound field including Spatial anchor, Commentary, Dialogue, and Object signals.
  • the metadata addition unit 12 also outputs the produced metadata to the coding unit 13.
  • the coding unit 13 Based on the mixed sound signals received from the mixing unit 11 and the metadata received from the metadata addition unit 12, the coding unit 13 produces the sound signals according to the Extended sound field descriptor, encodes the produced sound signals, and outputs the encoded sound signals to the multiplexer 14.
  • the multiplexer 14 receives, from the coding unit 13, the sound signals according to the Extended sound field descriptor that have been encoded, and multiplexes the received sound signals into a bit stream, in order to convey a multiplexed sound signal to a sound signal reproduction equipment via broadcast or transmission.
  • the multiplexer 14 transmits the multiplexed bit stream to remote places such as home via radio waves, IP circuits, and the like.
  • the monitoring unit 15 is used for checking contents of the sound signals and the metadata.
  • FIG. 3 shows a block diagram of the sound signal reproduction equipment according to one embodiment of the present invention.
  • the sound signal reproduction equipment utilizes the metadata included in the received sound signal and reproduces the received sound signal by controlling narration sound to be adjusted to a narration language and narration reproduction position desired by a user, while maintaining high quality sound providing as much of a sense of presence as was produced.
  • the sound signal reproduction equipment controls a sound image field position in the sound field layer of a "video/sound linked sound source", which requires a link between video and sound image positions, to be adjusted to the video display, and reproduces sound appropriately for reproduction environment with the video display, while maintaining the high quality sound providing as much of the sense of presence as was produced.
  • the sound signal reproduction equipment includes a demultiplexer 21, a decoding unit 22, a rendering reproduction unit 23, an environment information input unit 24, and monitoring unit 25.
  • the demultiplexer 21 receives, via broadcast or transmission, the sound signal according to the Extended sound field descriptor that has been multiplexed into the bit stream, and demultiplexes the received sound signal into the respective sound signals of the sound field layers and the metadata.
  • the demultiplexer 21 also outputs the demultiplexed sound signals and metadata to the decoding unit 22.
  • the decoding unit 22 decodes the encoded sound signals and metadata received from the demultiplexer 21 and outputs, to the rendering reproduction unit 23, signals including Spatial anchor, Commentary, Dialogue, Object signals, and metadata.
  • the rendering reproduction unit 23 Based on the Extended sound field descriptor, the rendering reproduction unit 23 reproduces the original sound signals as they are, or renders (e.g. down-mixes) the sound signals based on the reproduction environment (e.g. the number of channels of a speaker and a display size) before reproducing the sound signals. That is to say, the rendering reproduction unit 23 renders (e.g switches, converts, and renders) the sound signals based on the Extended sound field descriptor in a sound reproduction environment different from the environment during program production.
  • the reproduction environment e.g. the number of channels of a speaker and a display size
  • the environment information input unit 24 displays to a user the metadata information described as the Extended sound field descriptor, receives user inputs about the reproduction environment and demand information of the user, namely, language selection for the multiplexed sound, reproduction environment information (e.g. the speaker configuration and the display size), and the like, and outputs the input information to the rendering reproduction unit 23.
  • the metadata information described as the Extended sound field descriptor receives user inputs about the reproduction environment and demand information of the user, namely, language selection for the multiplexed sound, reproduction environment information (e.g. the speaker configuration and the display size), and the like, and outputs the input information to the rendering reproduction unit 23.
  • the monitoring unit 25 is used for checking a result of reproduction performed by the rendering reproduction unit 23, as well as program viewing.
  • the sound signal production equipment and the sound signal reproduction equipment according to the present invention make it possible to easily control the narration language switching and narration reproduction position relocation in accordance with the home reproduction environment and user demand.
  • the sound signal production equipment and the sound signal reproduction equipment according to the present invention make it possible to easily control the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the video to be linked to the sound image position, to be adjusted to the video display and perform reproduction, while maintaining the high quality sound providing as much of the sense of presence as was produced.
  • the metadata addition unit 12 adds the metadata shown in Table 10 to the header of the corresponding multichannel-sound-format signal or to the headers on each sound channel constituting the multichannel according to the Extended sound field descriptor.
  • the user inputs the information of the reproduction system, such as the speaker arrangement information and the user demand of narration sound position to be reproduced, and controls the sound signals (e.g. the user arbitrarily adjusts the reproduction position).
  • the sound signals can be reproduced under control in terms of a desired narration language and narration reproduction position while the high quality sound providing as much of the sense of presence as was produced is maintained.
  • the user at an receiving side inputs, through the environment information input unit 24, the information of desired narration sound (e.g. the narration language that the user demands to reproduce and the narration reproduction position) and the information of the reproduction system (e.g. speaker arrangement information).
  • the rendering reproduction unit 23 switches a sound signal of the "narration language” layer that has been designated from among the produced narration languages described in the metadata, adds to the switched sound signal the international sound used irrespective of language for reproduction, and reproduces the sound signal.
  • the rendering reproduction unit 23 is also fed the desired narration reproduction position, the speaker arrangement information, and the sound signal of the produced "narration language” layer.
  • the rendering reproduction unit 23 also relocates the switched sound signal so that reproduction is performed from the designated narration reproduction position and renders the signal so that the sound quality providing as much of the sense of presence as was produced is achieved. Subsequently, the rendering reproduction unit 23 adds, to the rendered signal, the international sound used irrespective of language and reproduces the signal.
  • FIG. 4 is a conceptual diagram of the multi-layered sound field including the sound field layer of the international sound (Spatial anchor) used irrespective of language, and the sound field layers of the "narration languages” (Commentary, Dialogue).
  • the sound signal production system is formed by the format of the "sound signals to compose a multi-layered sound field" including the sound field layer of the "sound requiring the link between video and sound positions” and the "sound directly irrespective of the video position.”
  • the metadata addition unit 12 adds the metadata shown in Table 11 to the header of the corresponding multichannel sound format signal or to the headers on each sound channel constituting the multichannel according to the Extended sound field descriptor.
  • the sound signal reproduction equipment controls the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the link between video and sound image positions, to be adjusted to the video display and reproduces sound, while maintaining the high quality sound providing as much of the sense of presence as was produced.
  • the user at the receiving side inputs, through the environment information input unit 24, the information of the reproduction system (e.g. speaker arrangement and video display information).
  • the rendering reproduction unit 23 does neither convert nor render the received sound signals. In this case, the rendering reproduction unit 23 adds the "sound requiring the link between video and sound positions" and the "sound directly irrespective of the video position" and reproduces the added sound.
  • the rendering reproduction unit 23 converts the received sound signals by either rendering or down-mixing so that the sound quality providing as much of the sense of presence as was produced is achieved, and reproduces the added sound signals.
  • the rendering reproduction unit 23 renders the sound signals of the layer of the "sound preferably requiring the link between video and sound positions" so that a width of the video display size equals a width of the sound image.
  • the rendering reproduction unit 23 adds the rendered "sound preferably requiring the link between video and sound positions" and the unconverted and un-rendered "sound directly irrespective of the video position" and reproduces the added sound.
  • the rendering processing i.e., processing for equalizing the width between the sound image of the "sound preferably requiring the link between video and sound positions" and the video display size, can be easily performed by using field position information of Azimuth angle and Elevation angle included in Spatial position data defined in Channel position data.
  • FIG. 6 is a conceptual diagram of the multi-layered sound field including the sound field layer of "video/sound linked sound source” (Video linked object) and the sound field layers “directly irrespective of the video position" (Spatial anchor, Dialogue).
  • the Extended sound field descriptor includes the number of sound field layers, the type of each sound field layer, and the language information.
  • the type of each sound field layer indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language.
  • the Extended sound field descriptor includes the number of multiple sound field layers and a video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video.
  • the sound signal described by the Extended sound field descriptor can be produced and reproduced.
  • the present invention also includes, in its scope, any equipment that transmits the sound signal described by the Extended sound field descriptor to the remote places such as home via radio waves, IP circuits, and the like, any equipment that stores and records in a recording medium the sound signal described by the Extended sound field descriptor, and a recording medium in which the sound signal described by the Extended sound field descriptor is stored and recorded.
  • the sound signal production equipment produces the metadata including the number of sound field layers, the type of each sound field layer, and the language information, produces the sound signal according to the Extended sound field descriptor based on an input sound signal and the metadata, and multiplexes the sound signal into the bit stream. Furthermore, the sound signal reproduction equipment according to one embodiment of the present invention converts the sound signal according to the number of sound field layers, the type of each sound field layer, and the language information included in the sound signal and according to the reproduction environment and demand information of the user, and reproduces the converted sound signal.
  • the above structure makes it possible to produce and view a program using the "sound signals to compose a multi-layered sound field."
  • the sound signal reproduction equipment adds, to the international sound, the sound signal of the particular language that has been switched by the user, and reproduces the added sound.
  • the above structure allows the user to arbitrarily carry out an operation such as language selection with use of the received metadata, thereby making it possible to switch and relocate the appropriate narration language and narration reproduction position, while the high quality sound providing as much of the sense of presence as was produced is maintained.
  • the sound signal production equipment produces the metadata including the number of layers of sound field and a video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video, produces the sound signal according to the Extended sound field descriptor based on the input sound signal and the metadata, and multiplexes the sound signal into the bit stream.
  • the sound signal reproduction equipment converts the sound signal according to the video link identifier and according to the reproduction environment information of the user, the video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video, and the sound signal reproduction equipment reproduces the converted sound signal.
  • the above structure makes it possible to produce and view the program using the "sound signals to compose a multi-layered sound field.”
  • the rendering reproduction unit renders the sound signal of the sound field layer based on information about the video display of the user, and reproduces the rendered sound signal.
  • the above structure makes it possible to render and convert the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the link between video and sound image positions, so that the sound field image position is adjusted to the video display, while the high quality sound providing as much of the sense of presence as was produced is maintained by inputting the information of the reproduction system (e.g. the video display) of the user and by using the information of the video display during production described in the metadata.
  • the present invention makes it possible to describe a "sound signals to compose a multi-layered sound field", and to produce and view/listen a program using such sound signals. As a result, interoperability between different next generation sound systems is achieved, and even in a sound reproduction environment different from the environment during program production, switching, conversion, and rendering of the sound signals is facilitated.
  • mixing unit 12 metadata addition unit 13
  • coding unit 14 multiplexer 15
  • monitoring unit 21 demultiplexer 22
  • decoding unit 23 rendering reproduction unit 24 environment information input unit 25 monitoring unit

Landscapes

  • Physics & Mathematics (AREA)
  • Engineering & Computer Science (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Stereophonic System (AREA)

Abstract

Provided is a sound signal description method corresponding to a format of "sound signals to compose a multi-layered sound field", as well as a sound signal production equipment and a sound signal reception equipment which correspond to the sound signal description method. The sound signal description method for describing the multi-layered sound field includes the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information.

Description

SOUND SIGNAL DESCRIPTION METHOD, SOUND SIGNAL PRODUCTION EQUIPMENT, AND SOUND SIGNAL REPRODUCTION EQUIPMENT
The present invention relates to a sound signal description method, a sound signal production equipment, and a sound signal reproduction equipment, all of which are capable of representing information of sound signals with use of metadata for sound reproduction through multichannel speakers.
Various sound systems, such as a 2 channel sound system, a 5.1 channel sound system, and "3-dimensional multichannel stereophonic sound systems" beyond the 5.1 channel sound system, are used for program production. Describing the various sound systems using a common description format provides flexibility to the sound systems, which allows the systems to be applied to next-generation sound systems across various sound application scenarios. ITU-R, which is an international standardization body associated with broadcasting including sound, has defined requirements for an advanced multichannel sound system as ITU-R Recommendation. (Refer to Non Patent Literature 1.)
[NPL 1] "Performance requirements for an advanced multichannel stereophonic sound system for use with or without accompanying picture", Recommendation ITU-R BS.1909.
As the common description format for describing the various sound systems, an advanced study has been conducted on "sound signals to compose a single-layered sound field." However, in some cases of sound program production, the format of "sound signals to compose a multi-layered sound field" can be used so as to facilitate rendering, conversion, and switching of received sound signals according to a receiver's environment or demand of program exchange or a home reproduction. For example, the receiver of program exchange or the home sometimes does not employ the same size image display as in the program production, and according to such a video reproduction environment of the receiver, the sound signal needs to be converted. Furthermore, it is sometimes required a language switching for program reproduction and, a reproduction position relocation of a narration signal according to needs of the receiver. Conventionally, however, the study has not been conducted on the description method for the "sound signals to compose a multi-layered sound field."
The present invention has been conceived in view of the above problems and aims to provide a sound signal description method corresponding to the format of the "sound signals to compose a multi-layered sound field", as well as a sound signal production equipment and a sound signal reproduction equipment which correspond to the sound signal description method.
In order to solve the aforementioned problems, one aspect of the present invention provides a sound signal description method for describing a multi-layered sound field, comprising: the number of sound field layers of the multi-layered sound field; a type of each sound field layer of the multi-layered sound field; and language information.
It is preferable that the type of each sound field layer of the multi-layered sound field indicates the sound elements of the program, such as one of international sound, which consists of all the sound program elements except for the commentary/dialogue elements, and one of commentary/dialogue sound with particular language.
Furthermore, another aspect of the present invention provides a sound signal description method for describing a multi-layered sound field, comprising: the number of sound field layers of the multi-layered sound field; and a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
Moreover, yet another aspect of the present invention provides a sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information; a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and a multiplexer that multiplexes the produced sound signal into a bit stream.
Moreover, yet another aspect of the present invention provides a sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: an environment information input unit that inputs reproduction environment and user demand information; and a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information included in the sound signal and according to the reproduction environment and user demand information, and reproduces the converted sound signal.
The type of each sound field layer of the multi-layered sound field indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language, and the particular language being switched by the environment information input unit. The rendering reproduction unit preferably adds the sound signal of the particular language to the international sound and reproduces added sound.
Moreover, yet another aspect of the present invention provides a sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field and a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video; a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and a multiplexer that multiplexes the produced sound signal into a bit stream.
Moreover, yet another aspect of the present invention provides a sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising: an environment information input unit that inputs reproduction environment information into the reproduction equipment; and a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field and a video link identifier included in the sound signal and according to the reproduction environment information. The video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
When the video link identifier indicates that the sound field layer is linked to video, the rendering reproduction unit preferably renders the sound signal of the sound field layer based on video display information input by the environment information input unit.
The sound signal description method, the sound signal production equipment, and the sound signal reproduction equipment according to the present invention make it possible to describe the "sound signals to compose a multi-layered sound field" and to produce and reproduce a sound program using the sound signals.
FIG. 1 shows an exemplary structure of an "Extended sound field descriptor" according to one embodiment of the present invention. FIG. 2 shows a block diagram of a sound signal production equipment according to one embodiment of the present invention. FIG. 3 shows a block diagram of a sound signal reproduction equipment according to one embodiment of the present invention. FIG. 4 is a conceptual diagram of a multi-layered sound field in connection with narration language switching. FIG. 5 shows a difference in display size between a program production environment and a reproduction environment. FIG. 6 is a conceptual diagram of the multi-layered sound field associated with linked/unlinked video and sound. FIG. 7 shows an exemplary structure of a "Basic sound field descriptor".
The following describes embodiments of the present invention in detail with reference to the drawings.
The present invention extends a description method (referred to below as a "Basic sound field descriptor") for describing "sound signals to compose a single-layered sound field" to the description method (referred to below as an "Extended sound field descriptor") for describing a "sound signals to compose a multi-layered sound field." Regarding the Basic sound field descriptor, the present applicants filed a Korean Patent Application (10-2012-0112984), and the Basic sound field descriptor is reviewed below for understanding of the present invention.
In order to describe multichannel sound signals to compose a single-layered sound field, it is necessary to describe which channel corresponds to the reproduction position. The described information is called descriptor, which is described as metadata in a header of a corresponding multichannel sound signal or in the headers on each sound channel constituting the multichannel.
Table 1 illustrates terms and definitions of the Basic sound field descriptor. The Basic sound field descriptor is employed for production and exchange of complete mix programs (i.e. programs including all sound required for reproduction) with multichannel sound, for example.
Figure JPOXMLDOC01-appb-T000001
The Sound Essence descriptor includes a descriptor of a program, a descriptor (name) of the Sound-field, and other relevant descriptors.
As shown in FIG. 7, the Sound-field is described by the Sound-field configuration with a hierarchical structure.
The Sound Channel descriptor includes the Channel label descriptor and/or Channel Position descriptor.
The following describes the descriptors in the Basic sound field descriptor. Note that some of the descriptors overlap with each other in anticipation of different program exchange scenarios. However, a program producer or the like is able to appropriately choose necessary descriptors for each program exchange scenario.
The Basic sound field descriptor includes (A) Sound Essence descriptors, (B) Sound-field configuration descriptors, and (C) Sound Channel descriptors.
Table 2 shows (A) Sound Essence descriptors in the Basic sound field descriptor.
Figure JPOXMLDOC01-appb-T000002
Table 3 shows (B) Sound-field configuration descriptors in the Basic sound field descriptor.
Figure JPOXMLDOC01-appb-T000003
Table 4 shows (C) Sound Channel descriptors in the Basic sound field descriptor.
Figure JPOXMLDOC01-appb-T000004
Table 5 shows C.1 Channel label descriptors, which are descriptors of the Channel label data included in the Sound Channel descriptors.
Figure JPOXMLDOC01-appb-T000005
Table 6 shows C.2 Channel position descriptors, which are descriptors of the Channel position data included in the Sound Channel descriptors.
Figure JPOXMLDOC01-appb-T000006
The present invention extends the Basic sound field descriptor, which is the description method for the "sound signals to compose a single-layered sound field" as mentioned above, to the Extended sound field descriptor, which is the description method for the "sound signals to compose a multi-layered sound field."
Table 7 illustrates terms and definitions of the Extended sound field descriptor.
Figure JPOXMLDOC01-appb-T000007
The Sound Essence descriptor includes the descriptor of the program, the descriptor (name) of the Sound-field, and the other relevant descriptors.
As shown in FIG. 1, the Sound-field in the Extended sound field descriptor is described by multiple Sound-field configurations (Group of sound-field configurations) (Sound space configurations) each having the hierarchical structure.
The Sound Channel descriptor includes the Channel label descriptor and/or the Channel Position descriptor.
Table 8 shows (A) Sound Essence descriptors in the Extended sound field descriptor.
Figure JPOXMLDOC01-appb-T000008
Table 9 shows A.2 Sound-field descriptors in the Extended sound field descriptor.
Figure JPOXMLDOC01-appb-T000009
Regarding (B) Sound-field configuration descriptors and (C) Sound Channel descriptors in the Extended sound field descriptor, these descriptors are the same as those of the Basic sound field descriptor, and a description thereof is omitted.
FIG. 2 shows a block diagram of a sound signal production equipment according to one embodiment of the present invention. In order to "facilitate" rendering, conversion, and switching of received sound signals according to the receiver's environment or demand of program exchange or the home reproduction, the sound signal production equipment produces a sound program according to the Extended sound field descriptor, which is the format of the "sound signals to compose a multi-layered sound field." The sound signal production equipment inserts the Extended sound field descriptor as metadata into the header of the corresponding sound format signal or into the header of each audio signal, for program exchange and transmission to the home. The sound signal production equipment includes a mixing unit 11, a metadata addition unit 12, a coding unit 13, a multiplexer 14, and a monitoring unit 15.
The mixing unit 11 mixes sound signals (Sound Sources 1-M) and outputs, to the coding unit 13, sound signals to compose the multi-layered sound field including Spatial anchor, Commentary, Dialogue, and Object signals, the sound signals being output from a "production system for sound signals to compose a multi-layered sound field."
The metadata addition unit 12 outputs, to the coding unit 13, the metadata to be described for the Extended sound field descriptor of the multi-layered sound field including Spatial anchor, Commentary, Dialogue, and Object signals. The metadata addition unit 12 also outputs the produced metadata to the coding unit 13.
Based on the mixed sound signals received from the mixing unit 11 and the metadata received from the metadata addition unit 12, the coding unit 13 produces the sound signals according to the Extended sound field descriptor, encodes the produced sound signals, and outputs the encoded sound signals to the multiplexer 14.
The multiplexer 14 receives, from the coding unit 13, the sound signals according to the Extended sound field descriptor that have been encoded, and multiplexes the received sound signals into a bit stream, in order to convey a multiplexed sound signal to a sound signal reproduction equipment via broadcast or transmission. The multiplexer 14 transmits the multiplexed bit stream to remote places such as home via radio waves, IP circuits, and the like.
The monitoring unit 15 is used for checking contents of the sound signals and the metadata.
FIG. 3 shows a block diagram of the sound signal reproduction equipment according to one embodiment of the present invention. In accordance with an input of information about a reproduction system, such as speaker arrangement information and user demand of narration sound position to be reproduced, the sound signal reproduction equipment utilizes the metadata included in the received sound signal and reproduces the received sound signal by controlling narration sound to be adjusted to a narration language and narration reproduction position desired by a user, while maintaining high quality sound providing as much of a sense of presence as was produced. Furthermore, in a reproduction environment with a video display having a different size from a size according to production conditions, the sound signal reproduction equipment controls a sound image field position in the sound field layer of a "video/sound linked sound source", which requires a link between video and sound image positions, to be adjusted to the video display, and reproduces sound appropriately for reproduction environment with the video display, while maintaining the high quality sound providing as much of the sense of presence as was produced. The sound signal reproduction equipment includes a demultiplexer 21, a decoding unit 22, a rendering reproduction unit 23, an environment information input unit 24, and monitoring unit 25.
The demultiplexer 21 receives, via broadcast or transmission, the sound signal according to the Extended sound field descriptor that has been multiplexed into the bit stream, and demultiplexes the received sound signal into the respective sound signals of the sound field layers and the metadata. The demultiplexer 21 also outputs the demultiplexed sound signals and metadata to the decoding unit 22.
The decoding unit 22 decodes the encoded sound signals and metadata received from the demultiplexer 21 and outputs, to the rendering reproduction unit 23, signals including Spatial anchor, Commentary, Dialogue, Object signals, and metadata.
Based on the Extended sound field descriptor, the rendering reproduction unit 23 reproduces the original sound signals as they are, or renders (e.g. down-mixes) the sound signals based on the reproduction environment (e.g. the number of channels of a speaker and a display size) before reproducing the sound signals. That is to say, the rendering reproduction unit 23 renders (e.g switches, converts, and renders) the sound signals based on the Extended sound field descriptor in a sound reproduction environment different from the environment during program production.
The environment information input unit 24 displays to a user the metadata information described as the Extended sound field descriptor, receives user inputs about the reproduction environment and demand information of the user, namely, language selection for the multiplexed sound, reproduction environment information (e.g. the speaker configuration and the display size), and the like, and outputs the input information to the rendering reproduction unit 23.
The monitoring unit 25 is used for checking a result of reproduction performed by the rendering reproduction unit 23, as well as program viewing.
The following describes specific usage embodiments of the sound signal production equipment and the sound signal reproduction equipment. For example, the sound signal production equipment and the sound signal reproduction equipment according to the present invention make it possible to easily control the narration language switching and narration reproduction position relocation in accordance with the home reproduction environment and user demand. Furthermore, in the reproduction environment with the video display having the different size than the size according to production conditions, the sound signal production equipment and the sound signal reproduction equipment according to the present invention make it possible to easily control the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the video to be linked to the sound image position, to be adjusted to the video display and perform reproduction, while maintaining the high quality sound providing as much of the sense of presence as was produced.
(Production Embodiment 1: Production of Signal Including Sound Field Layer Associated with multiple Languages)
As an example of program production using the Extended sound field descriptor, i.e., the format of the "sound signals to compose a multi-layered sound field", suppose a case where not only the sound signals of the Japanese or Korean narrations and dialogues but also the sound signals of various languages such as English are produced. In the above example, the sound signal production system is formed by the format of the "sound signals to compose a multi-layered sound field" including the sound field layer of the international sound (Spatial anchor) used irrespective of language, and the sound field layers (Commentary, Dialogue) of the narrations and dialogues of particular languages.
In this example, the metadata addition unit 12 adds the metadata shown in Table 10 to the header of the corresponding multichannel-sound-format signal or to the headers on each sound channel constituting the multichannel according to the Extended sound field descriptor.
Figure JPOXMLDOC01-appb-T000010
(Reproduction Embodiment 1: Reproduction of Signal Including Sound Field Layer Associated with Multiple Languages)
The user inputs the information of the reproduction system, such as the speaker arrangement information and the user demand of narration sound position to be reproduced, and controls the sound signals (e.g. the user arbitrarily adjusts the reproduction position). For example, in the home reproduction environment the sound signals can be reproduced under control in terms of a desired narration language and narration reproduction position while the high quality sound providing as much of the sense of presence as was produced is maintained.
In order to achieve the above function, the user at an receiving side inputs, through the environment information input unit 24, the information of desired narration sound (e.g. the narration language that the user demands to reproduce and the narration reproduction position) and the information of the reproduction system (e.g. speaker arrangement information). The rendering reproduction unit 23 switches a sound signal of the "narration language" layer that has been designated from among the produced narration languages described in the metadata, adds to the switched sound signal the international sound used irrespective of language for reproduction, and reproduces the sound signal. The rendering reproduction unit 23 is also fed the desired narration reproduction position, the speaker arrangement information, and the sound signal of the produced "narration language" layer. The rendering reproduction unit 23 also relocates the switched sound signal so that reproduction is performed from the designated narration reproduction position and renders the signal so that the sound quality providing as much of the sense of presence as was produced is achieved. Subsequently, the rendering reproduction unit 23 adds, to the rendered signal, the international sound used irrespective of language and reproduces the signal.
FIG. 4 is a conceptual diagram of the multi-layered sound field including the sound field layer of the international sound (Spatial anchor) used irrespective of language, and the sound field layers of the "narration languages" (Commentary, Dialogue).
(Production Embodiment 2: Production of Program Including Sound Field Layer Associated with Linked/Unlinked Video and Sound)
As an example of program production using the Extended sound field descriptor, i.e., the format of the "sound signals to compose a multi-layered sound field", suppose a case where the "sound requiring the link between video and sound positions" and the "sound directly irrespective of the video position" are separately produced and recorded. Sound signals include not only the "sound requiring the link between video and sound positions" (e.g. the dialogue of an actor and sound emitted from an object on the screen) but also the "sound directly irrespective of the video position" (e.g. sound effects for enhancing the sense of presence of an entire program), and the "sound requiring the link between video and sound positions" and the "sound directly irrespective of the video position" can be separately produced and recorded. In the above example, the sound signal production system is formed by the format of the "sound signals to compose a multi-layered sound field" including the sound field layer of the "sound requiring the link between video and sound positions" and the "sound directly irrespective of the video position."
In this example, the metadata addition unit 12 adds the metadata shown in Table 11 to the header of the corresponding multichannel sound format signal or to the headers on each sound channel constituting the multichannel according to the Extended sound field descriptor.
Figure JPOXMLDOC01-appb-T000011
(Reproduction Embodiment 2: Reproduction of Program Including Sound Field Layer Associated with Linked/Unlinked Video and Sound)
In the reproduction environment with the video display having the different size than the size according to the production conditions as shown in FIG. 5, for example, the sound signal reproduction equipment controls the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the link between video and sound image positions, to be adjusted to the video display and reproduces sound, while maintaining the high quality sound providing as much of the sense of presence as was produced.
In order to achieve the above function, the user at the receiving side inputs, through the environment information input unit 24, the information of the reproduction system (e.g. speaker arrangement and video display information). When the conditions for the video display and the speaker arrangement during production are the same as the conditions for the video display and the speaker arrangement at the receiving side, the rendering reproduction unit 23 does neither convert nor render the received sound signals. In this case, the rendering reproduction unit 23 adds the "sound requiring the link between video and sound positions" and the "sound directly irrespective of the video position" and reproduces the added sound. On the other hand, when the above conditions are not the same in terms of either one of the video display and the speaker arrangement, the rendering reproduction unit 23 converts the received sound signals by either rendering or down-mixing so that the sound quality providing as much of the sense of presence as was produced is achieved, and reproduces the added sound signals. When the video display size is different, and the speaker arrangement is the same, the rendering reproduction unit 23 renders the sound signals of the layer of the "sound preferably requiring the link between video and sound positions" so that a width of the video display size equals a width of the sound image. The rendering reproduction unit 23 adds the rendered "sound preferably requiring the link between video and sound positions" and the unconverted and un-rendered "sound directly irrespective of the video position" and reproduces the added sound. Here, the rendering processing, i.e., processing for equalizing the width between the sound image of the "sound preferably requiring the link between video and sound positions" and the video display size, can be easily performed by using field position information of Azimuth angle and Elevation angle included in Spatial position data defined in Channel position data.
FIG. 6 is a conceptual diagram of the multi-layered sound field including the sound field layer of "video/sound linked sound source" (Video linked object) and the sound field layers "directly irrespective of the video position" (Spatial anchor, Dialogue).
Thus, according to the above embodiment, the Extended sound field descriptor includes the number of sound field layers, the type of each sound field layer, and the language information. With the above structure, the sound signal description method corresponding to the format of the "sound signals to compose a multi-layered sound field" is achieved.
Furthermore, it is preferable that the type of each sound field layer indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language. With the above structure, in the home reproduction environment, for example, the sound signals can be reproduced under control in terms of the desired narration language and narration reproduction position while the high quality sound providing as much of the sense of presence as was produced is maintained.
Moreover, according to the above embodiment, the Extended sound field descriptor includes the number of multiple sound field layers and a video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video. With the above structure, in the reproduction environment with the video display having the different size than the size according to the production conditions, for example, the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the link between video and sound image positions, can be controlled to be adjusted to the video display, and reproduction is performed, while the high quality sound providing as much of the sense of presence as was produced is maintained.
Moreover, with the sound signal production equipment and the sound signal reproduction equipment according to the above embodiments, the sound signal described by the Extended sound field descriptor can be produced and reproduced. Note that the present invention also includes, in its scope, any equipment that transmits the sound signal described by the Extended sound field descriptor to the remote places such as home via radio waves, IP circuits, and the like, any equipment that stores and records in a recording medium the sound signal described by the Extended sound field descriptor, and a recording medium in which the sound signal described by the Extended sound field descriptor is stored and recorded.
The sound signal production equipment according to one embodiment of the present invention produces the metadata including the number of sound field layers, the type of each sound field layer, and the language information, produces the sound signal according to the Extended sound field descriptor based on an input sound signal and the metadata, and multiplexes the sound signal into the bit stream. Furthermore, the sound signal reproduction equipment according to one embodiment of the present invention converts the sound signal according to the number of sound field layers, the type of each sound field layer, and the language information included in the sound signal and according to the reproduction environment and demand information of the user, and reproduces the converted sound signal. The above structure makes it possible to produce and view a program using the "sound signals to compose a multi-layered sound field." In particular, the sound signal reproduction equipment adds, to the international sound, the sound signal of the particular language that has been switched by the user, and reproduces the added sound. The above structure allows the user to arbitrarily carry out an operation such as language selection with use of the received metadata, thereby making it possible to switch and relocate the appropriate narration language and narration reproduction position, while the high quality sound providing as much of the sense of presence as was produced is maintained.
Moreover, the sound signal production equipment according to one embodiment of the present invention produces the metadata including the number of layers of sound field and a video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video, produces the sound signal according to the Extended sound field descriptor based on the input sound signal and the metadata, and multiplexes the sound signal into the bit stream. Moreover, the sound signal reproduction equipment according to one embodiment of the present invention converts the sound signal according to the video link identifier and according to the reproduction environment information of the user, the video link identifier indicating, for each sound field layer, whether the sound field layer is linked to video, and the sound signal reproduction equipment reproduces the converted sound signal. The above structure makes it possible to produce and view the program using the "sound signals to compose a multi-layered sound field." In particular, when the video link identifier indicates that the sound field layer is linked to video, the rendering reproduction unit renders the sound signal of the sound field layer based on information about the video display of the user, and reproduces the rendered sound signal. The above structure makes it possible to render and convert the sound image field position in the sound field layer of the "video/sound linked sound source", which requires the link between video and sound image positions, so that the sound field image position is adjusted to the video display, while the high quality sound providing as much of the sense of presence as was produced is maintained by inputting the information of the reproduction system (e.g. the video display) of the user and by using the information of the video display during production described in the metadata.
While the present invention has been described based on the figures and embodiments, it should be noted that a person skilled in the art can readily make various modifications and changes in accordance with the disclosure. As such, it should also be noted that the modifications and changes are within the scope of the present invention. For example, the function or the like included in each element, each means, and each step is subject to rearrangement, and several means and steps can be combined into a single means or step or they can be divided.
The present invention makes it possible to describe a "sound signals to compose a multi-layered sound field", and to produce and view/listen a program using such sound signals. As a result, interoperability between different next generation sound systems is achieved, and even in a sound reproduction environment different from the environment during program production, switching, conversion, and rendering of the sound signals is facilitated.
11 mixing unit
12 metadata addition unit
13 coding unit
14 multiplexer
15 monitoring unit
21 demultiplexer
22 decoding unit
23 rendering reproduction unit
24 environment information input unit
25 monitoring unit

Claims (9)

  1. A sound signal description method for describing a multi-layered sound field, comprising:
    the number of sound field layers of the multi-layered sound field;
    a type of each sound field layer of the multi-layered sound field; and
    language information.
  2. The sound signal description method recited in claim 1, wherein the type of each sound field layer of the multi-layered sound field indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language.
  3. A sound signal description method for describing a multi-layered sound field, comprising:
    the number of sound field layers of the multi-layered sound field; and
    a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
  4. A sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising:
    a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information;
    a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and
    a multiplexer that multiplexes the produced sound signal into a bit stream.
  5. A sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising:
    an environment information input unit that inputs reproduction environment and user demand information; and
    a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field, a type of each sound field layer of the multi-layered sound field, and language information included in the sound signal and according to the reproduction environment and user demand information, and reproduces the converted sound signal.
  6. The sound signal reproduction equipment recited in claim 5, wherein
    the type of each sound field layer of the multi-layered sound field indicates which one of international sound and a particular language the sound field layer comprises, the international sound being used irrespective of language, and the particular language being switched by the environment information input unit, and
    the rendering reproduction unit adds the sound signal of the particular language to the international sound and reproduces added sound.
  7. A sound signal production equipment that produces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising:
    a metadata addition unit that produces metadata including the number of sound field layers of the multi-layered sound field and a video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video;
    a coding unit that produces the sound signal according to the sound signal description method based on an input sound signal and the metadata; and
    a multiplexer that multiplexes the produced sound signal into a bit stream.
  8. A sound signal reproduction equipment that reproduces a sound signal according to a sound signal description method for describing a multi-layered sound field, comprising:
    an environment information input unit that inputs reproduction environment and user demand information; and
    a rendering reproduction unit that converts the sound signal according to the number of sound field layers of the multi-layered sound field and a video link identifier included in the sound signal and according to the reproduction and user demand environment information, the video link identifier indicating, for each sound field layer of the multi-layered sound field, whether the sound field layer is linked to video.
  9. The sound signal reproduction equipment recited in claim 8, wherein
    when the video link identifier indicates that the sound field layer is linked to video, the rendering reproduction unit renders the sound signal of the sound field layer based on video display information input by the environment information input unit.
PCT/JP2013/007390 2013-01-23 2013-12-16 Sound signal description method, sound signal production equipment, and sound signal reproduction equipment Ceased WO2014115222A1 (en)

Priority Applications (2)

Application Number Priority Date Filing Date Title
US14/652,907 US20150334502A1 (en) 2013-01-23 2013-12-16 Sound signal description method, sound signal production equipment, and sound signal reproduction equipment
KR1020157018270A KR101682323B1 (en) 2013-01-23 2013-12-16 Sound signal description method, sound signal production equipment, and sound signal reproduction equipment

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2013-010544 2013-01-23
JP2013010544A JP6174326B2 (en) 2013-01-23 2013-01-23 Acoustic signal generating device and acoustic signal reproducing device

Publications (1)

Publication Number Publication Date
WO2014115222A1 true WO2014115222A1 (en) 2014-07-31

Family

ID=51227039

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2013/007390 Ceased WO2014115222A1 (en) 2013-01-23 2013-12-16 Sound signal description method, sound signal production equipment, and sound signal reproduction equipment

Country Status (4)

Country Link
US (1) US20150334502A1 (en)
JP (1) JP6174326B2 (en)
KR (1) KR101682323B1 (en)
WO (1) WO2014115222A1 (en)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10109285B2 (en) 2014-09-08 2018-10-23 Sony Corporation Coding device and method, decoding device and method, and program

Families Citing this family (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
EP4601259A3 (en) 2014-09-30 2025-09-24 Sony Group Corporation Transmitting device, transmission method, receiving device, and receiving method
US10032447B1 (en) * 2014-11-06 2018-07-24 John Mitchell Kochanczyk System and method for manipulating audio data in view of corresponding visual data
WO2018173413A1 (en) * 2017-03-24 2018-09-27 シャープ株式会社 Audio signal processing device and audio signal processing system
CN109286888B (en) * 2018-10-29 2021-01-29 中国传媒大学 Audio and video online detection and virtual sound image generation method and device
KR102658473B1 (en) 2021-03-17 2024-04-18 한국전자통신연구원 METHOD AND APPARATUS FOR Label encoding in polyphonic sound event intervals
KR102914075B1 (en) 2021-04-20 2026-01-19 한국전자통신연구원 Method and system for processing obstacle effect in virtual acoustic space
JP2022188830A (en) * 2021-06-10 2022-12-22 日本放送協会 Object-based acoustic coordinate transform device and program

Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07264144A (en) * 1994-03-16 1995-10-13 Toshiba Corp Signal compression coding apparatus and compressed signal decoding apparatus
JPH0950673A (en) * 1995-08-02 1997-02-18 Sony Corp Data encoding apparatus and method, data recording medium, and data decoding apparatus and method
JP2005049436A (en) * 2003-07-30 2005-02-24 Toshiba Corp Speech recognition method, apparatus and program
JP2008154065A (en) * 2006-12-19 2008-07-03 Roland Corp Effect imparting device
JP2009278381A (en) * 2008-05-14 2009-11-26 Nippon Hoso Kyokai <Nhk> Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information
JP2010521013A (en) * 2007-03-09 2010-06-17 エルジー エレクトロニクス インコーポレイティド Audio signal processing method and apparatus
JP2011035784A (en) * 2009-08-04 2011-02-17 Sharp Corp Stereoscopic video-stereophonic sound recording and reproducing device, system, and method
JP2011234177A (en) * 2010-04-28 2011-11-17 Panasonic Corp Stereoscopic sound reproduction device and reproduction method

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7844355B2 (en) * 2005-02-18 2010-11-30 Panasonic Corporation Stream reproduction device and stream supply device
CN101026725B (en) * 2005-07-15 2010-09-29 索尼株式会社 Reproducing apparatus and reproducing method
JP4897404B2 (en) * 2006-09-12 2012-03-14 株式会社ソニー・コンピュータエンタテインメント VIDEO DISPLAY SYSTEM, VIDEO DISPLAY DEVICE, ITS CONTROL METHOD, AND PROGRAM
GB2442983A (en) * 2006-10-16 2008-04-23 Martin Audio Ltd A computer-based method of configuring loudspeaker arrays
KR20090001605A (en) * 2007-05-03 2009-01-09 삼성전자주식회사 Apparatus and method for reproducing content by using a reproducing setting information recorded on a removable recording medium and a reproducing setting information
US8908874B2 (en) * 2010-09-08 2014-12-09 Dts, Inc. Spatial audio encoding and reproduction
CN103299616A (en) * 2011-01-07 2013-09-11 夏普株式会社 Reproducing device and control method thereof, generating device and control method thereof, recording medium, data structure, control program, and recording medium recording the program
PL2727381T3 (en) * 2011-07-01 2022-05-02 Dolby Laboratories Licensing Corporation Apparatus and method for rendering audio objects
BR112015013154B1 (en) * 2012-12-04 2022-04-26 Samsung Electronics Co., Ltd Audio delivery device, and audio delivery method

Patent Citations (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH07264144A (en) * 1994-03-16 1995-10-13 Toshiba Corp Signal compression coding apparatus and compressed signal decoding apparatus
JPH0950673A (en) * 1995-08-02 1997-02-18 Sony Corp Data encoding apparatus and method, data recording medium, and data decoding apparatus and method
JP2005049436A (en) * 2003-07-30 2005-02-24 Toshiba Corp Speech recognition method, apparatus and program
JP2008154065A (en) * 2006-12-19 2008-07-03 Roland Corp Effect imparting device
JP2010521013A (en) * 2007-03-09 2010-06-17 エルジー エレクトロニクス インコーポレイティド Audio signal processing method and apparatus
JP2009278381A (en) * 2008-05-14 2009-11-26 Nippon Hoso Kyokai <Nhk> Acoustic signal multiplex transmission system, manufacturing device, and reproduction device added with sound image localization acoustic meta-information
JP2011035784A (en) * 2009-08-04 2011-02-17 Sharp Corp Stereoscopic video-stereophonic sound recording and reproducing device, system, and method
JP2011234177A (en) * 2010-04-28 2011-11-17 Panasonic Corp Stereoscopic sound reproduction device and reproduction method

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10109285B2 (en) 2014-09-08 2018-10-23 Sony Corporation Coding device and method, decoding device and method, and program
US10446160B2 (en) 2014-09-08 2019-10-15 Sony Corporation Coding device and method, decoding device and method, and program

Also Published As

Publication number Publication date
JP2014142475A (en) 2014-08-07
US20150334502A1 (en) 2015-11-19
JP6174326B2 (en) 2017-08-02
KR101682323B1 (en) 2016-12-02
KR20150093794A (en) 2015-08-18

Similar Documents

Publication Publication Date Title
WO2014115222A1 (en) Sound signal description method, sound signal production equipment, and sound signal reproduction equipment
JP5174527B2 (en) Acoustic signal multiplex transmission system, production apparatus and reproduction apparatus to which sound image localization acoustic meta information is added
US10771175B2 (en) Receiving apparatus, receiving method, transmitting apparatus, and transmitting method
CN114930876B (en) Systems, methods and apparatus for conversion from channel-based audio to object-based audio
EP2094032A1 (en) Audio signal, method and apparatus for encoding or transmitting the same and method and apparatus for processing the same
JP6407155B2 (en) Audio data generating apparatus and audio data reproducing apparatus
JP2013174891A (en) High quality multi-channel audio encoding and decoding apparatus
JP2011150358A (en) Method and apparatus for processing two or more audio signals
KR20090115074A (en) Method and apparatus for transmitting and receiving multichannel audio signal using super frame
US11039212B2 (en) Reception device
JP6204681B2 (en) Acoustic signal reproduction device
Mróz et al. A commonly-accessible toolchain for live streaming music events with higher-order ambisonic audio and 4k 360 vision
JP6204682B2 (en) Acoustic signal reproduction device
Grewe et al. MPEG-H Audio production workflows for a Next Generation Audio Experience in Broadcast, Streaming and Music
JP2014204322A (en) Acoustic signal reproducing device and acoustic signal preparation device
JP6204684B2 (en) Acoustic signal reproduction device
JP6670802B2 (en) Sound signal reproduction device
JP6204680B2 (en) Acoustic signal reproduction device, acoustic signal creation device
Grewe et al. Mpeg-h audio system for sbtvd tv 3.0 call for proposals
Bolt et al. Practical implementation of new open standards for Next Generation Audio production and interchange
JP2014222856A (en) Acoustic signal reproduction device and acoustic signal preparation device
JP2014204316A (en) Acoustic signal reproducing device and acoustic signal preparation device
KR101346709B1 (en) Server for transmitting a multiview video, apparatus for receiving a multiview video and system and method for multiview broadcast using thereof
Poers Metadata based audio production for Next Generation Audio formats
Weitnauer et al. Promising Noises About Media Sound? A Progress Report on the NGA/AdvSS Developments in the ITU-R in 2018/2019

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 13873093

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 14652907

Country of ref document: US

ENP Entry into the national phase

Ref document number: 20157018270

Country of ref document: KR

Kind code of ref document: A

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 13873093

Country of ref document: EP

Kind code of ref document: A1