WO2018107372A1 - 一种声音处理方法、装置、电子设备及计算机程序产品 - Google Patents

一种声音处理方法、装置、电子设备及计算机程序产品 Download PDF

Info

Publication number
WO2018107372A1
WO2018107372A1 PCT/CN2016/109774 CN2016109774W WO2018107372A1 WO 2018107372 A1 WO2018107372 A1 WO 2018107372A1 CN 2016109774 W CN2016109774 W CN 2016109774W WO 2018107372 A1 WO2018107372 A1 WO 2018107372A1
Authority
WO
WIPO (PCT)
Prior art keywords
real
orientation
right ear
human body
audio
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2016/109774
Other languages
English (en)
French (fr)
Inventor
骆磊
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Cloudminds Shenzhen Robotics Systems Co Ltd
Original Assignee
Cloudminds Shenzhen Robotics Systems Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Cloudminds Shenzhen Robotics Systems Co Ltd filed Critical Cloudminds Shenzhen Robotics Systems Co Ltd
Priority to CN201680002674.2A priority Critical patent/CN107077318A/zh
Priority to PCT/CN2016/109774 priority patent/WO2018107372A1/zh
Publication of WO2018107372A1 publication Critical patent/WO2018107372A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/16Sound input; Sound output
    • G06F3/162Interface to dedicated audio devices, e.g. audio drivers, interface to CODECs
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F30/00Computer-aided design [CAD]
    • G06F30/20Design optimisation, verification or simulation

Definitions

  • Embodiments of the present invention relate to the field of audio processing technologies, and in particular, to a sound processing method, apparatus, electronic device, and computer program product.
  • 3D games and virtual reality/augmented reality can realize 360-degree visual presentation, and the mixture of real and virtual brings the user experience to a new height visually.
  • the embodiment of the invention provides a sound processing method, device, electronic device and computer program product, and the main technical problem is to improve the authenticity of the output sound.
  • a technical solution adopted by the embodiment of the present invention is to provide a sound processing method, including: determining a left ear position and a right ear position of a human body model in a real-time scene; and simulating an audio to be played in the real-time scene.
  • the sound at the left ear position is output as a left ear audio, simulating the sound effect of the audio to be played in the real-time scene at the right ear position as the right ear audio output.
  • the determining the left ear position and the right ear position of the human body model in the real-time scene includes: determining the position and orientation of the human body model; determining the left ear position and the right ear position of the human body model according to the position and orientation of the human body model.
  • determining the left ear position and the right ear position of the human body model according to the position and orientation comprising: determining a head position and a head orientation according to the position and orientation of the human body model; A human head model having the orientation of the human head is implanted at a corresponding head position; a left ear position on the human head model is acquired as a left ear position of the human body model, and a right ear position on the human head model is acquired as a right ear position of the human body model.
  • the sound effect of the audio to be played in the simulated real-time scene at the position of the left ear includes: according to the material of each object in the real-time scene and the orientation of the utterance point relative to the left ear, the audio corresponding to the position of the corresponding utterance point is performed.
  • Processing the sound effect at the position of the left ear; the sound effect of the audio to be played in the simulated real-time scene at the position of the right ear includes: according to the material of each object in the real-time scene and the orientation of the utterance point relative to the right ear, The audio corresponding to the position of the corresponding sound point is processed to obtain the sound effect at the position of the right ear.
  • the scene is a game scene.
  • a sound processing device including: a positioning module, configured to determine a left ear and a right ear position of a human body model in a real-time scene; The sound effect of the audio to be played in the simulated real-time scene at the left ear position is output as the left ear audio, and the sound effect of the audio to be played in the real-time scene at the right ear position is simulated as the right ear audio output.
  • the positioning module is configured to determine the position and orientation of the human body model, and determine the left ear position and the right ear position of the human body model according to the position and orientation of the human body model.
  • the positioning module is configured to determine a head position and a head orientation according to the position and orientation of the human body model, implant a human head model having the head orientation on the corresponding head position, and obtain a left ear position on the head model as The left ear position of the mannequin is obtained as the right ear position of the human body model as the right ear position on the human head model.
  • the simulation module is configured to process the audio corresponding to the position of the corresponding sound point according to the material of each object in the real-time scene and the orientation of the sound point relative to the left ear, and obtain the sound effect at the left ear position;
  • the material of each object in the real-time scene and the orientation of the utterance point relative to the right ear, and the audio corresponding to the position of the corresponding utterance point is processed to obtain the sound effect at the position of the right ear.
  • an electronic device including: at least one processor; a memory, the memory is connected to a processor, and the memory is stored At least one processor executing a computer executable A line instruction, the computer executable instructions being executed by the at least one processor to cause the at least one processor to perform the sound processing method described above.
  • the electronic device is an augmented reality device or a virtual reality device.
  • another technical solution adopted by the embodiment of the present invention is to provide a computer program product, the computer program product comprising a computing program stored on a non-transitory computer readable storage medium, the computer
  • the program includes program instructions that, when executed by the computer, cause the calculation to perform the above-described sound processing method.
  • the beneficial effects of the embodiments of the present invention are: different from the prior art, the sound processing method, the device, the electronic device, and the computer program product in the embodiment of the present invention pass the information of the position and sound of each original sounding point in real time.
  • the sound information of the recording position of the HATS model in the simulation space is transmitted to the earphone to obtain the most realistic panoramic sound.
  • FIG. 1 is a schematic flow chart of a sound processing method according to an embodiment of the present invention.
  • FIG. 2 is a schematic diagram of a state in which a sound processing method is applied to an augmented reality game scene according to an embodiment of the present invention
  • FIG. 3 is a schematic diagram of determining the positions of the left ear and the right ear of the human body model in the scenario of FIG. 2;
  • FIG. 4 is a schematic flow chart of a method for determining a left ear position and a right ear position of a human body model according to the position and orientation of a human body model in a sound processing method according to an embodiment of the present invention
  • FIG. 5 is a schematic structural diagram of a sound processing apparatus according to an embodiment of the present invention.
  • FIG. 6 is a schematic structural diagram of an electronic device for implementing a sound processing method according to an embodiment of the present invention.
  • FIG. 1 is a schematic flowchart of a sound processing method according to an embodiment of the present invention.
  • the method shown in this embodiment includes:
  • step S10 the left ear position and the right ear position of the human body model in the real-time scene are determined.
  • the scene is a game scene.
  • the left ear position and the right ear position of the human body model in the real-time scene are determined, specifically: determining the position and orientation of the human body model, and determining the left ear position and the right ear position of the human body model according to the position and orientation of the human body model.
  • the position of the binaural of the human torso model in the scene three-dimensional model is set as a simulation point, and the three-dimensional coordinate values of the sound objects in the three-dimensional coordinate system of the scene in the scene are obtained in real time. And the audio information, as well as the three-dimensional coordinate values and angles of the current time simulation point in the three-dimensional coordinate system of the scene three-dimensional model.
  • the user's position is replaced by a virtual head-rotatable HATS model (Head and Torso Simulator) (which coincides with the user's upper body).
  • HATS head is centered in three dimensions.
  • the real-time position in the stereo map is marked as (0,0,0).
  • the character is always in the center of the picture, and the coordinates of all other positions in the map are transformed in real time during the walking process.
  • the ability of the head to rotate can be different. For example, in a traditional 3D game, the person's head can only rotate up and down, and cannot swing or rotate left and right (the left and right rotation is rotated with the body). In the latest VR/AR scene, the head can be rotated in three axes.
  • determining the left ear position and the right ear position of the human body model according to the position and orientation of the human body model include:
  • step S31 the head position and the head orientation are determined according to the position and orientation of the mannequin.
  • step S32 a human head model having the orientation of the human head is implanted at the corresponding head position.
  • step S33 the left ear position on the human head model is acquired as the left ear position of the human body model, and the right ear position on the human head model is acquired as the right ear position of the human body model.
  • the user position is replaced by the HATS model
  • the coordinates of the center point of the head and the ear are set to (0, 0, 0)
  • the direction of the trunk is rotated by a degree of Z (false)
  • Set the positive direction of the trunk to coincide with the positive direction of the X-axis and record it as 0 degrees. It is assumed that the top direction coincides with the Z axis, and the front of the face coincides with the X axis as the positive direction of the head.
  • the coordinates of the left and right ears are (0, d/2, 0) and (0, -d/2, 0), respectively, and the face faces the positive direction of the X axis. , recorded as 0 degrees.
  • Step S11 simulating the sound effect of the audio to be played in the real-time scene at the left ear position as the left ear audio output, and simulating the sound effect of the audio to be played in the real-time scene at the right ear position as the right ear audio output.
  • the sound effect of the audio that needs to be played in the real-time scene at the position of the left ear is specifically: according to the material of each object in the real-time scene and the orientation of the utterance point relative to the left ear, the audio corresponding to the position of the corresponding utterance point is processed. Get the sound at the left ear position.
  • the sound effect of the audio that needs to be played in the real-time scene at the position of the right ear is specifically: according to the material of each object in the real-time scene and the orientation of the utterance point relative to the right ear, the audio corresponding to the position of the corresponding utterance point is processed. Get the sound at the right ear position.
  • a head-rotatable HATS model is simulated in real time in the position and direction of the character, and the position of the headphones placed on the left and right ears of the HATS model is taken in real time according to the way of recording the human body trunk model. point.
  • the real-time three-dimensional scene model information is used to input the information such as the sound position, position and direction of each sound point, the position, direction and head rotation angle of the HATS model, and the audio information of the recording record point of the HATS model is obtained through simulation. And output this audio directly to the headphones worn by the user, you can get the most realistic panoramic sound experience.
  • the sound information of the head position of the head of the HTS model that can be rotated in the space is calculated in real time and transmitted to the earphone to obtain the most realistic panoramic sound.
  • the sound processing method of the embodiment of the present invention is explained by taking the scenario as an example.
  • the scene is a virtual reality (VR) based game scene.
  • VR virtual reality
  • the user may have various sounds such as footsteps, voices, gunshots, explosions, etc. in different locations around the scene, because the information is controlled by the game itself, so the specific coordinate position, direction, The specific sound of the source is known.
  • the position of the binaural model of the human torso model in the scene three-dimensional model is set as a simulation point, and the three-dimensional coordinate value and audio information of each sound object in the three-dimensional coordinate system of the scene in the scene are acquired in real time, and the current time simulation point is obtained.
  • the three-dimensional coordinate values and angles in the three-dimensional coordinate system of the scene three-dimensional model are set as a simulation point, and the three-dimensional coordinate value and audio information of each sound object in the three-dimensional coordinate system of the scene in the scene are acquired in real time, and the current time simulation point is obtained.
  • the user's position is replaced by a virtual head rotatable HATS human torso model (coincident with the user's upper body), and the real position of the HATS head's ear center in the three-dimensional map at any time. Marked as (0,0,0). The character is always in the center of the picture, and the coordinates of all other positions in the map are transformed in real time during the walking process.
  • the real-time position of the center of the head and ears in the three-dimensional map is (0, 0, 0), and the position of a grenade on the side of the ground is (3, 5, -1.7), and when the character is kneeling, The real-time position of the center of both ears of the head in the three-dimensional map is still (0,0,0), and the position of the handel may be changed to (3,5,-0.6).
  • the character's armpit causes the Z-axis position to change, and the position of the person's head and the ground grenades is close on the Z-axis).
  • the X and Y axes are horizontal planes, and Z is the height axis perpendicular to the horizontal plane:
  • the user position is replaced by the HATS model.
  • the coordinates of the center point of the head and the ear are set to (0, 0, 0), and the direction of the trunk is a degree rotated by the Z axis (assuming that the positive direction of the trunk is directly forward and coincides with the positive direction of the X axis. Recorded as 0 degrees).
  • the top direction coincides with the Z axis
  • the front of the face coincides with the X axis as the positive direction of the head.
  • the distance between the ears is d
  • the coordinates of the left and right ears are (0, d/2, 0) and (0, -d/2, 0), respectively, and the face faces the positive direction of the X axis. , recorded as 0 degrees.
  • the angles of the two ears and the X, Y and Z axes are ⁇ , ⁇ , ⁇ at a certain time.
  • the coordinates of the two ears are (cos ⁇ *d/2, cos ⁇ *d/2, cos ⁇ *d/2) and (cos( ⁇ – ⁇ )*d/2, cos( ⁇ – ⁇ )*d, respectively.
  • /2, cos( ⁇ – ⁇ )*d/2) the left and right ears are always axisymmetric.
  • the angle ⁇ , ⁇ , ⁇ , and the orientation of the face with the three axes can be directly obtained by the motion sensor passing through the current game or experience device (such as a g-sensor of a mobile phone, a gyroscope, a geomagnetic sensor).
  • the motion sensor passing through the current game or experience device such as a g-sensor of a mobile phone, a gyroscope, a geomagnetic sensor.
  • the three-dimensional coordinates in the scene map, the sound, coordinates, and direction of each utterance point are also It is known that the left and right ear positions (cos ⁇ *d/2, cos ⁇ *d/2, cos ⁇ *d/2) and (cos( ⁇ – ⁇ )*d/2, cos( ⁇ – ⁇ )*d are obtained in real time. /2, cos( ⁇ - ⁇ ) * d / 2) simulation sound information of the simulation point.
  • the simulation results can be transmitted to the headphones worn by the user in real time to get the most realistic results.
  • the audio information of the sound object is processed, and according to the audio information of the processed sound object, the orientation of the sound object in the three-dimensional model of the scene, and the information of the simulation point, for each sound
  • the information of the object is encoded to generate corresponding three-dimensional audio. Therefore, according to the material of each surface in the scene of the three-dimensional model of the scene, more realistic simulations of the reverberation, damping and other changes and the head-rotatable HATS body torso model, all the information that affects the sound propagation and change has been determined.
  • the left and right ear positions (cos ⁇ *d/2, cos ⁇ *d/2, cos ⁇ *d/2) and (cos( ⁇ – ⁇ )*d/2,cos( ⁇ – ⁇ )*d/2 can be calculated in real time.
  • FIG. 5 is a schematic structural diagram of a sound processing apparatus according to an embodiment of the present invention.
  • the device 50 includes a positioning module 51 and an analog module 52.
  • the positioning module 51 is configured to determine the left ear and right ear positions of the human body model in the real-time scene.
  • the simulation module 52 is configured to simulate the sound effect of the audio to be played in the real-time scene at the left ear position as the left ear audio output, and simulate the sound effect of the audio to be played in the real-time scene at the right ear position as the right ear audio output.
  • the positioning module 51 is configured to determine the position and orientation of the human body model, and determine the left ear position and the right ear position of the human body model according to the position and orientation of the human body model.
  • the positioning module 51 is configured to determine a head position and a head orientation according to the position and orientation of the mannequin, implant a head model with a head orientation on the corresponding head position, and acquire a left ear position on the head model as a human body model.
  • the left ear position is obtained as the right ear position on the human head model as the right ear position of the human body model.
  • the simulation module 52 is configured to process the audio corresponding to the position of the corresponding sound point according to the material of each object in the real-time scene and the orientation of the sound point relative to the left ear, and obtain the sound effect at the left ear position;
  • the material of each object and the orientation of the sounding point relative to the right ear, and the audio corresponding to the position of the corresponding sounding point is processed to obtain the position of the right ear.
  • the sound effect is configured to process the audio corresponding to the position of the corresponding sound point according to the material of each object in the real-time scene and the orientation of the sound point relative to the left ear, and obtain the sound effect at the left ear position;
  • the material of each object and the orientation of the sounding point relative to the right ear, and the audio corresponding to the position of the corresponding sounding point is processed to obtain the position of the right ear.
  • the sound effect is configured to process the audio corresponding to the position of the corresponding sound point according to the material of each object in the real-time scene
  • FIG. 6 is a schematic structural diagram of an electronic device according to an embodiment of the present invention.
  • the electronic device 60 in this embodiment may be a computer.
  • the electronic device 60 includes a receiver 61, a processor 62, a transmitter 63, a read only memory 64, a random access memory 65, and a bus 66.
  • the receiver 61 is for receiving data.
  • the processor 62 can also be a CPU (Central Processing Unit).
  • the processor 62 may be an integrated circuit chip with signal processing capabilities.
  • Processor 62 can also be a general purpose processor, digital signal processor (DSP), application specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic device, discrete hardware component .
  • DSP digital signal processor
  • ASIC application specific integrated circuit
  • FPGA off-the-shelf programmable gate array
  • the general purpose processor may be a microprocessor or the processor or any conventional processor or the like.
  • the transmitter 63 is for transmitting data.
  • the memory can include read only memory 64 and random access memory 65 and provide instructions and data to processor 62.
  • a portion of the memory may also include non-volatile random access memory (NVRAM).
  • NVRAM non-volatile random access memory
  • bus 66 which may include, in addition to the data bus, a power bus, a control bus, a status signal bus, and the like. However, for clarity of description, various buses are labeled as bus 66 in the figure.
  • the memory stores the following elements, executable modules or data structures, or a subset of them, or their extended set:
  • Operation instructions include various operation instructions for implementing various operations.
  • Operating system Includes a variety of system programs for implementing various basic services and handling hardware-based tasks.
  • the processor 62 performs the following operations by calling an operation instruction stored in the memory, which can be stored in the operating system:
  • the sound effect of the audio to be played in the real-time scene at the left ear position is output as the left ear audio output
  • the sound effect of the audio to be played in the real-time scene at the right ear position is simulated as the right ear audio output.
  • the electronic device is an augmented reality device or a virtual reality device.
  • a computer program product comprising a computing program stored on a non-transitory computer readable storage medium, the computer program comprising program instructions, when the program instructions are executed by the computer, causing the computing computer to perform the following operations:
  • the sound effect of the audio to be played in the real-time scene at the left ear position is output as the left ear audio output
  • the sound effect of the audio to be played in the real-time scene at the right ear position is simulated as the right ear audio output.
  • the sound processing method, device, electronic device and computer program product of the embodiment of the present invention simulate the sound information of the recording position of the HATS model in the space in real time through the information of the position and sound of each original sounding point, and transmit the sound information to the earphone to Get the most realistic panoramic sound. Get a true panoramic sound.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Health & Medical Sciences (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • Multimedia (AREA)
  • Computer Hardware Design (AREA)
  • Evolutionary Computation (AREA)
  • Geometry (AREA)
  • Stereophonic System (AREA)

Abstract

一种声音处理方法、装置、电子设备及计算机程序产品。其中,该方法包括:确定实时场景中人体模型的左耳位置和右耳位置(S10);模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出(S11)。能够实现实时仿真场景中人体模型的左耳位置和右耳位置的声音信息,并传输给耳机,以得到最真实的全景声。

Description

一种声音处理方法、装置、电子设备及计算机程序产品 技术领域
本发明实施例涉及音频处理技术领域,特别是涉及一种声音处理方法、装置、电子设备及计算机程序产品。
背景技术
目前3D游戏以及虚拟现实/增强现实(VR/AR)能够实现360度的视觉呈现,真实与虚拟的混合,从视觉上将用户体验带到了一个新的高度。
然而,虽然视觉上做到了近乎真实的全景感受,听觉上离真实却还有明显差距,在头部转动的时候,声音没有明显变化,或者虽然有变化却感觉并不真实,更像模拟出来的3D音效的体验,定位感较差。
发明内容
本发明实施例提供一种声音处理方法、装置、电子设备及计算机程序产品,主要解决的技术问题是提高输出的声音的真实性。
为解决上述技术问题,本发明实施例采用的一个技术方案是:提供一种声音处理方法,包括:确定实时场景中人体模型的左耳位置和右耳位置;模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
其中,所述确定实时场景中人体模型的左耳位置和右耳位置,包括:确定人体模型的位置以及朝向;根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
其中,所述根据所述位置和朝向确定人体模型的左耳位置和右耳位置,包括:根据人体模型的位置以及朝向确定人头位置和人头朝向;在 相应的人头位置上植入一个具有所述人头朝向的人头模型;获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
其中,所述模拟实时场景中需要播放的音频在左耳位置处的音效,包括:根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效;所述模拟实时场景中需要播放的音频在右耳位置处的音效,包括:根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置处的音效。
其中,所述场景为游戏场景。
为解决上述技术问题,本发明实施例采用的另一个技术方案是:提供一种声音处理装置,包括:定位模块,用于确定实时场景中人体模型的左耳和右耳位置;模拟模块,用于模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
其中,所述定位模块用于确定人体模型的位置以及朝向,并根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
其中,所述定位模块用于根据人体模型的位置以及朝向确定人头位置和人头朝向,在相应的人头位置上植入一个具有所述人头朝向的人头模型,以及获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
其中,所述模拟模块用于根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效;还用于根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置处的音效。
为解决上述技术问题,本发明实施例采用的另一个技术方案是:提供一种电子设备,包括:至少一个处理器;存储器,所述存储器与处理器连接,所述存储器存储有可被所述至少一个处理器执行的计算机可执 行指令,所述计算机可执行指令被所述至少一个处理器执行,以使所述至少一个处理器用于执行上述的声音处理方法。
其中,所述电子设备为增强现实设备或者虚拟现实设备。
为解决上述技术问题,本发明实施例采用的另一个技术方案是:提供一种计算机程序产品,所述计算机程序产品包括存储在非易失性计算机可读存储介质上的计算程序,所述计算机程序包括程序指令,当所述程序指令被计算机执行时,使所述计算执机行上述的声音处理方法。
本发明实施例的有益效果是:区别于现有技术的情况,本发明实施例中的声音处理方法、装置、电子设备及计算机程序产品,通过每一个原始发声点的位置和声音等信息,实时仿真空间中HATS模型录音点位置的声音信息,并传输给耳机,以得到最真实的全景声。
附图说明
图1是本发明实施例中一种声音处理方法的流程示意图;
图2是本发明实施例中的一种声音处理方法应用于增强现实游戏场景的状态示意图;
图3是在图2的场景下确定人体模型的左耳、右耳位置的示意图;
图4是本发明实施例中的一种声音处理方法中根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置的方法流程示意图;
图5是本发明实施例中的一种声音处理装置的结构示意图;
图6是本发明实施例中的一种实现种声音处理方法的电子设备的结构示意图。
具体实施方式
为了便于理解本发明,下面结合附图和具体实施例,对本发明进行更详细的说明。
除非另有定义,本说明书所使用的所有的技术和科学术语与属于本发明的技术领域的技术人员通常理解的含义相同。本说明书中在本发明的说明书中所使用的术语只是为了描述具体的实施例的目的,不是用于限制本发明。本说明书所使用的术语“和/或”包括一个或多个相关的所列 项目的任意的和所有的组合。
请参阅图1,是本发明实施例中一种声音处理方法的流程示意图。该实施例示出的方法包括:
步骤S10,确定实时场景中人体模型的左耳位置和右耳位置。
其中,该场景为游戏场景。
确定实时场景中人体模型的左耳位置和右耳位置,具体为:确定人体模型的位置以及朝向,并根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
请同时参阅图2,具体地,将人体躯干模型的双耳在场景三维模型中的位置设置为仿真点,实时获取该场景中各声音对象在该场景三维模型的三维坐标系中的三维坐标值和音频信息,以及当前时刻仿真点在场景三维模型的三维坐标系中的三维坐标值和角度。
在该场景地图中将用户的位置换成一个虚拟的头部可转动的HATS模型(Head and Torso Simulator,人头与躯干模型)(与用户上身重合),任何时候将HATS头部两耳中心在三维立体地图中的实时位置标记为(0,0,0)。人物始终在画面中心,走动过程中实时变换地图中所有其他位置的坐标。根据使用场景的不同,头部可转动的能力也会不同,如传统3D游戏中,人的头部只能上下转动,不能够摆动或者左右转动(左右转动是和身体一起转动)。而最新的VR/AR场景中,头部可以以三轴进行转动。
进一步地,请参阅图4,根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置,包括:
步骤S31,根据人体模型的位置以及朝向确定人头位置和人头朝向。
步骤S32,在相应的人头位置上植入一个具有所述人头朝向的人头模型。
步骤S33,获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
请再次参阅图2,具体地,用户位置用HATS模型替换,头部两耳连线中心点坐标设置为(0,0,0),躯干方向为以Z为轴转动a度(假 设躯干正方向为正前方与X轴正方向重合,记为0度。假设以头顶方向与Z轴重合,并且面部正前方与X轴重合为头部的正方向。两耳间距为d,则可知在头部正方向的情况下,左右耳坐标分别为(0,d/2,0)和(0,-d/2,0),脸部面向X轴正方向,记为0度。
步骤S11,模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
其中,模拟实时场景中需要播放的音频在左耳位置处的音效,具体为:根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效。
其中,模拟实时场景中需要播放的音频在右耳位置处的音效,具体为:根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置处的音效。
具体地,根据场景中的三维立体模型信息,在人物所处位置和方向实时模拟一个头部可转动的HATS模型,按照人体躯干模型录音的方式,取HATS模型左右耳放置耳机的位置为实时仿真点。以实时的三维立体场景模型信息,将各个发声点声音、位置、方向等信息、HATS模型的位置、方向、头部转动角度等信息作为输入,通过仿真计算得到HATS模型录音记录点的音频信息,并将此音频直接输出给用户佩戴的耳机,即可得到最真实的全景声音体验。
以上内容,通过记录每一个原始发声点的声音、位置、方向,实时仿真计算空间中头部可转动的HATS模型人头录双耳位置的声音信息,并传输给耳机,以得到最真实的全景声。
请再次参阅图2,下面以该场景为例对本发明实施例的声音处理方法进行解释说明。在本实施例中,该场景为基于虚拟现实(VR)射击游戏场景。
在某一特定时刻,用户在该场景中,周围不同位置可能有脚步声、说话声、枪声、爆炸声等种种声音,因为这些信息都是由游戏本身控制的,所以具体坐标位置、方向、源头的具体声音都是已知的。
将人体躯干模型的双耳在场景三维模型中的位置设置为仿真点,实时获取该场景中各声音对象在该场景三维模型的三维坐标系中的三维坐标值和音频信息,以及当前时刻仿真点在场景三维模型的三维坐标系中的三维坐标值和角度。
具体地,在该场景地图中将用户的位置换成一个虚拟的头部可转动的HATS人体躯干模型(与用户上身重合),任何时候将HATS头部两耳中心在三维立体地图中的实时位置标记为(0,0,0)。人物始终在画面中心,走动过程中实时变换地图中所有其他位置的坐标。例如,人物站立时,头部两耳中心在三维立体地图中的实时位置为(0,0,0),旁边地上一个手雷位置为(3,5,-1.7),而当人物蹲下时,其头部两耳中心在三维立体地图中的实时位置依然为(0,0,0),而手雷位置可能变更为(3,5,-0.6)。(人物蹲下使得Z轴位置发生变化,人物头部与地上手雷的位置在Z轴上接近了)。
请再次参阅图2,仍以如上场景为例,X,Y轴为水平面,Z为垂直于水平面的高度轴:
用户位置用HATS模型替换,头部两耳连线中心点坐标设置为(0,0,0),躯干方向为以Z轴转动a度(假设躯干正方向为正前方与X轴正方向重合,记为0度)。
假设以头顶方向与Z轴重合,并且面部正前方与X轴重合为头部的正方向。两耳间距为d,则可知在头部正方向的情况下,左右耳坐标分别为(0,d/2,0)和(0,-d/2,0),脸部面向X轴正方向,记为0度。
请同时参阅图2、图3,假设某时刻两耳连线与X,Y,Z三轴的角度分别为α,β,γ。根据数学原理可知,两耳坐标分别为(cosα*d/2,cosβ*d/2,cosγ*d/2)和(cos(π–α)*d/2,cos(π–β)*d/2,cos(π–γ)*d/2),左右耳始终为轴对称。
如上,与三轴的夹角α,β,γ,以及面部朝向,躯干朝向通过当前游戏或体验设备的动作传感器即可直接获得(如手机的g-sensor,陀螺仪,地磁传感器)。
在该场景地图中的三维坐标,各个发声点的声音、坐标、方向也都 是已知,实时计算得到左右耳位置(cosα*d/2,cosβ*d/2,cosγ*d/2)和(cos(π–α)*d/2,cos(π–β)*d/2,cos(π–γ)*d/2)的仿真点的仿真声音信息。将此仿真结果实时传送到用户佩戴的耳机上即可得到最真实的效果。
进一步地,根据场景三维模型的表面材质,对声音对象的音频信息进行处理,以及根据处理后的声音对象的音频信息、声音对象在场景三维模型中的方位和仿真点的信息,对每个声音对象的信息进行编码以生成相应的三维音频。从而,根据场景三维模型场景中各表面的材质,以更真实的模拟出回响,阻尼等变化和头部可转动的HATS体躯干模型,所有影响声音传播和变化的信息都已经是确定的了,则可实时计算得到左右耳位置(cosα*d/2,cosβ*d/2,cosγ*d/2)和(cos(π–α)*d/2,cos(π–β)*d/2,cos(π–γ)*d/2)的仿真点的仿真声音信息。
请参阅图5,为本发明实施例中的一种声音处理装置的结构示意图。该装置50包括定位模块51和模拟模块52。
定位模块51用于确定实时场景中人体模型的左耳和右耳位置。
模拟模块52用于模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
进一步地,定位模块51用于确定人体模型的位置以及朝向,并根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
进一步地,定位模块51用于根据人体模型的位置以及朝向确定人头位置和人头朝向,在相应的人头位置上植入一个具有人头朝向的人头模型,以及获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
模拟模块52用于根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效;还用于根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置 处的音效。
请参阅图6,为本发明实施例中的电子设备的结构示意图,该实施例中的电子设备60可以是计算机。该电子设备60包括接收器61、处理器62、发送器63、只读存储器64、随机存取存储器65以及总线66。
该接收器61用于接收数据。
该处理器62还可以成为CPU(Central Processing Unit,中央处理单元)。该处理器62可能是一种集成电路芯片,具有信号的处理能力。处理器62还可以是通用处理器、数字信号处理器(DSP)、专用集成电路(ASIC)、现成可编程门阵列(FPGA)或者其他可编程逻辑器件、分立门或者晶体管逻辑器件、分立硬件组件。通用处理器可以是微处理器或者该处理器也可以是任何常规的处理器等。
该发送器63用于发送数据。
存储器可以包括只读存储器64和随机存取存储器65,并向处理器62提供指令和数据。存储器的一部分还可以包括非易失性随机存取存储器(NVRAM)。
电子设备60的各个组件通过总线66耦合在一起,其中,总线66除包括数据总线之外,还可以包括电源总线、控制总线和状态信号总线等。但是为了清楚说明起见,在图中将各种总线都标为总线66。
存储器存储了如下的元素,可执行模块或者数据结构,或者它们的子集,或者它们的扩展集:
操作指令:包括各种操作指令,用于实现各种操作。
操作系统:包括各种系统程序,用于实现各种基础业务以及处理基于硬件的任务。
在本发明实施例中,处理器62通过调用存储器存储的操作指令(该操作指令可存储在操作系统中),执行如下操作:
确定实时场景中人体模型的左耳位置和右耳位置;
模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
所述电子设备为增强现实设备或者虚拟现实设备。
一种计算机程序产品,计算机程序产品包括存储在非易失性计算机可读存储介质上的计算程序,计算机程序包括程序指令,当程序指令被计算机执行时,使计算执机行如下操作:
确定实时场景中人体模型的左耳位置和右耳位置;
模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
本发明实施例的声音处理方法、装置、电子设备及计算机程序产品,通过每一个原始发声点的位置和声音等信息,实时仿真空间中HATS模型录音点位置的声音信息,并传输给耳机,以得到最真实的全景声。得到真实全景声音。
需要说明的是,本发明的说明书及其附图中给出了本发明的较佳的实施例,但是,本发明可以通过许多不同的形式来实现,并不限于本说明书所描述的实施例,这些实施例不作为对本发明内容的额外限制,提供这些实施例的目的是使对本发明的公开内容的理解更加透彻全面。并且,上述各技术特征继续相互组合,形成未在上面列举的各种实施例,均视为本发明说明书记载的范围;进一步地,对本领域普通技术人员来说,可以根据上述说明加以改进或变换,而所有这些改进和变换都应属于本发明所附权利要求的保护范围。

Claims (12)

  1. 一种声音处理方法,其特征在于,包括:
    确定实时场景中人体模型的左耳位置和右耳位置;
    模拟实时场景中需要播放的音频在左耳位置处的音效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
  2. 根据权利要求1所述的方法,其特征在于,所述确定实时场景中人体模型的左耳位置和右耳位置,包括:
    确定人体模型的位置以及朝向;
    根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
  3. 根据权利要求2所述的方法,其特征在于,所述根据所述位置和朝向确定人体模型的左耳位置和右耳位置,包括:
    根据人体模型的位置以及朝向确定人头位置和人头朝向;
    在相应的人头位置上植入一个具有所述人头朝向的人头模型;
    获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
  4. 根据权利要求1所述的方法,其特征在于,所述模拟实时场景中需要播放的音频在左耳位置处的音效,包括:
    根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效;
    所述模拟实时场景中需要播放的音频在右耳位置处的音效,包括:
    根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置处的音效。
  5. 根据权利要求1所述的方法,其特征在于,所述场景为游戏场景。
  6. 一种声音处理装置,其特征在于,包括:
    定位模块,用于确定实时场景中人体模型的左耳和右耳位置;
    模拟模块,用于模拟实时场景中需要播放的音频在左耳位置处的音 效作为左耳音频输出,模拟实时场景中需要播放的音频在右耳位置处的音效作为右耳音频输出。
  7. 根据权利要求6所述的装置,其特征在于,所述定位模块用于确定人体模型的位置以及朝向,并根据人体模型的位置以及朝向确定人体模型的左耳位置和右耳位置。
  8. 根据权利要求7所述的装置,其特征在于,所述定位模块用于根据人体模型的位置以及朝向确定人头位置和人头朝向,在相应的人头位置上植入一个具有所述人头朝向的人头模型,以及获取人头模型上的左耳位置作为人体模型的左耳位置,获取人头模型上的右耳位置作为人体模型的右耳位置。
  9. 根据权利要求6所述的装置,其特征在于,所述模拟模块用于根据实时场景中的各个物体的材质以及发声点相对于左耳的方位,对相应发声点位置对应的音频进行处理得到在左耳位置处的音效;还用于根据实时场景中的各个物体的材质以及发声点相对于右耳的方位,对相应发声点位置对应的音频进行处理得到在右耳位置处的音效。
  10. 一种电子设备,其特征在于,包括:
    至少一个处理器;
    存储器,所述存储器与处理器连接,所述存储器存储有可被所述至少一个处理器执行的计算机可执行指令,所述计算机可执行指令被所述至少一个处理器执行,以使所述至少一个处理器用于执行如权利要求1-5任意一项所述的方法。
  11. 根据权利要求10所述的电子设备,其特征在于,所述电子设备为增强现实设备或者虚拟现实设备。
  12. 一种计算机程序产品,其特征在于,所述计算机程序产品包括存储在非易失性计算机可读存储介质上的计算程序,所述计算机程序包括程序指令,当所述程序指令被计算机执行时,使所述计算执机行如权利要求1-5中任意一项所述的方法。
PCT/CN2016/109774 2016-12-14 2016-12-14 一种声音处理方法、装置、电子设备及计算机程序产品 Ceased WO2018107372A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN201680002674.2A CN107077318A (zh) 2016-12-14 2016-12-14 一种声音处理方法、装置、电子设备及计算机程序产品
PCT/CN2016/109774 WO2018107372A1 (zh) 2016-12-14 2016-12-14 一种声音处理方法、装置、电子设备及计算机程序产品

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2016/109774 WO2018107372A1 (zh) 2016-12-14 2016-12-14 一种声音处理方法、装置、电子设备及计算机程序产品

Publications (1)

Publication Number Publication Date
WO2018107372A1 true WO2018107372A1 (zh) 2018-06-21

Family

ID=59623875

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2016/109774 Ceased WO2018107372A1 (zh) 2016-12-14 2016-12-14 一种声音处理方法、装置、电子设备及计算机程序产品

Country Status (2)

Country Link
CN (1) CN107077318A (zh)
WO (1) WO2018107372A1 (zh)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114356068A (zh) * 2020-09-28 2022-04-15 北京搜狗智能科技有限公司 一种数据处理方法、装置和电子设备

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2019067443A1 (en) 2017-09-27 2019-04-04 Zermatt Technologies Llc AUDIO SPACE NAVIGATION
CN111480348B (zh) * 2017-12-21 2022-01-07 脸谱公司 用于基于音频的增强现实的系统和方法
CN109107158B (zh) * 2018-09-04 2020-09-22 Oppo广东移动通信有限公司 音效处理方法、装置、电子设备及计算机可读存储介质
CN114777563B (zh) * 2022-05-20 2023-08-08 中国人民解放军总参谋部第六十研究所 一种应用于实兵对抗训练的战场声效模拟系统和方法

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2007112756A2 (en) * 2006-04-04 2007-10-11 Aalborg Universitet System and method tracking the position of a listener and transmitting binaural audio data to the listener
CN103702274A (zh) * 2013-12-27 2014-04-02 三星电子(中国)研发中心 立体环绕声重建方法及装置
CN104010265A (zh) * 2013-02-22 2014-08-27 杜比实验室特许公司 音频空间渲染设备及方法
CN104041081A (zh) * 2012-01-11 2014-09-10 索尼公司 声场控制装置、声场控制方法、程序、声场控制系统和服务器
CN105163242A (zh) * 2015-09-01 2015-12-16 深圳东方酷音信息技术有限公司 一种多角度3d声回放方法及装置
CN105872940A (zh) * 2016-06-08 2016-08-17 北京时代拓灵科技有限公司 一种虚拟现实声场生成方法及系统

Family Cites Families (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN105183421B (zh) * 2015-08-11 2018-09-28 中山大学 一种虚拟现实三维音效的实现方法及系统
CN106154231A (zh) * 2016-08-03 2016-11-23 厦门傅里叶电子有限公司 虚拟现实中声场定位的方法
CN106162206A (zh) * 2016-08-03 2016-11-23 北京疯景科技有限公司 全景录制、播放方法及装置

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2007112756A2 (en) * 2006-04-04 2007-10-11 Aalborg Universitet System and method tracking the position of a listener and transmitting binaural audio data to the listener
CN104041081A (zh) * 2012-01-11 2014-09-10 索尼公司 声场控制装置、声场控制方法、程序、声场控制系统和服务器
CN104010265A (zh) * 2013-02-22 2014-08-27 杜比实验室特许公司 音频空间渲染设备及方法
CN103702274A (zh) * 2013-12-27 2014-04-02 三星电子(中国)研发中心 立体环绕声重建方法及装置
CN105163242A (zh) * 2015-09-01 2015-12-16 深圳东方酷音信息技术有限公司 一种多角度3d声回放方法及装置
CN105872940A (zh) * 2016-06-08 2016-08-17 北京时代拓灵科技有限公司 一种虚拟现实声场生成方法及系统

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN114356068A (zh) * 2020-09-28 2022-04-15 北京搜狗智能科技有限公司 一种数据处理方法、装置和电子设备
CN114356068B (zh) * 2020-09-28 2023-08-25 北京搜狗智能科技有限公司 一种数据处理方法、装置和电子设备

Also Published As

Publication number Publication date
CN107077318A (zh) 2017-08-18

Similar Documents

Publication Publication Date Title
JP7839859B2 (ja) 双方向オーディオ環境のための空間オーディオ
US10453175B2 (en) Separate time-warping for a scene and an object for display of virtual reality content
JP7649798B2 (ja) 拡張現実インタラクションを提供するための装置および方法
TWI647593B (zh) 模擬環境顯示系統及方法
WO2018107372A1 (zh) 一种声音处理方法、装置、电子设备及计算机程序产品
WO2018196469A1 (zh) 声场的音频数据的处理方法及装置
CN110536665A (zh) 使用虚拟回声定位来仿真空间感知
CN106909335A (zh) 模拟声音源的方法
JP7714074B2 (ja) 反響利得正規化
US20250086908A1 (en) Device and method for providing augmented reality content
TWI731326B (zh) 高保真度環繞聲格式之音效處理系統及音效處理方法
WO2024131204A1 (zh) 虚拟场景设备交互方法及相关产品
US11915371B2 (en) Method and apparatus of constructing chess playing model
US20240406666A1 (en) Sound field capture with headpose compensation
JP2021527353A (ja) 低周波数チャネル間コヒーレンス制御
CN116764195A (zh) 基于虚拟现实vr的音频控制方法、装置、电子设备及介质
CN120929041A (zh) 音效配置方法、装置及电子设备
CN119421100A (zh) 音频处理方法、装置、电子设备及可读介质
CN120050590A (zh) 虚拟声源的音频信号分配

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 16923999

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 05/11/2019)

122 Ep: pct application non-entry in european phase

Ref document number: 16923999

Country of ref document: EP

Kind code of ref document: A1