WO2026005309A1 - System and method of enhanced audio rendering for an immersive virtual reality experience - Google Patents

System and method of enhanced audio rendering for an immersive virtual reality experience

Info

Publication number
WO2026005309A1
WO2026005309A1 PCT/KR2025/007342 KR2025007342W WO2026005309A1 WO 2026005309 A1 WO2026005309 A1 WO 2026005309A1 KR 2025007342 W KR2025007342 W KR 2025007342W WO 2026005309 A1 WO2026005309 A1 WO 2026005309A1
Authority
WO
WIPO (PCT)
Prior art keywords
audio
video frame
sound
virtual reality
spatial
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/KR2025/007342
Other languages
French (fr)
Inventor
Sandeep Singh SPALL
Choice CHOUDHARY
Ankit Agarwal
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Samsung Electronics Co Ltd
Original Assignee
Samsung Electronics Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Samsung Electronics Co Ltd filed Critical Samsung Electronics Co Ltd
Publication of WO2026005309A1 publication Critical patent/WO2026005309A1/en
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/305Electronic adaptation of stereophonic audio signals to reverberation of the listening space
    • H04S7/306For headphones
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04SSTEREOPHONIC SYSTEMS 
    • H04S7/00Indicating arrangements; Control arrangements, e.g. balance control
    • H04S7/30Control circuits for electronic adaptation of the sound field
    • H04S7/302Electronic adaptation of stereophonic sound system to listener position or orientation
    • H04S7/303Tracking of listener position or orientation
    • H04S7/304For headphones

Definitions

  • the present disclosure generally relates to the field of virtual reality systems, and more specifically relates to a system and method of an enhanced audio rendering for an immersive virtual reality experience to a user within a virtual environment.
  • Motion sickness in virtual reality (VR) environments is a common issue that many individuals experience. This phenomenon occurs when there is a disconnect between what the eyes see and what the inner ear senses, leading to feelings of nausea, dizziness, and discomfort. When users are exposed to rapid movements or visuals that do not align with their physical movements, the brain can become confused, resulting in the motion sickness.
  • Various techniques have been developed to mitigate the symptoms, primarily such solutions are based on reducing latency, adjusting a field of view, and creating smoother transitions between movements.
  • Low-latency image rendering to the current viewport is a standard solution for managing visual aspects of motion sickness.
  • such conventional methods often fail to consider the auditory component, which is crucial for maintaining an immersive AR/VR experience.
  • the eyes receive information from the virtual environment almost instantly, forming a perception of orientation, movement, and space.
  • the visual content is out of synchronization with the body’s movements, this leads to motion sickness.
  • the conventional systems fail to adjust audio with the same low latency and precision as video, causing a breakdown in the coordination between the image and body movement.
  • VR controllers may be used by the user to interact with the VR environment.
  • the problem arises due to inconsistencies between the rendered VR image at the user’s viewport and the audio rendering.
  • the user experiences audio that is unnatural, distorted, or unsynchronized, which may compromise the immersive experience.
  • a VR engine updates a visual scene to match the new viewpoint.
  • an audio engine may not update with the same speed and accuracy, leading to inconsistencies.
  • the sounds may become distorted or fail to match the visual cues accurately. This creates a jarring and unnatural experience for the user.
  • VR content may include unsynchronized audio that does not match the visual movements and actions within the VR environment.
  • the environment highlights the challenges faced in the conventional VR systems where the synchronization between the visual and auditory elements is not adequately managed. These inconsistencies result in an overall degraded VR experience, with the users experiencing unnatural, distorted, or unsynchronized audio, especially during head movements.
  • an electronic apparatus comprising: a memory, at least one processor comprising a processing circuit, wherein the at least one processor configured to obtain a content associated with a virtual reality environment, wherein the content includes audio data and image data, obtain video frame information associated with the image data, wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data, obtain audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics.
  • the audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
  • the at least one processor may output the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  • the sound spatial position may indicate a location of the sound source in three-dimensional space relative to a user's position.
  • the at least one processor may update the sound spatial position in real-time based on a change of the video frame speed.
  • the at least one processor may obtain the audio spatial characteristics based on one or more spatial audio cues in the audio data, wherein the one or more spatial audio cues may include at least one of increasing or decreasing speed associated with a sound source.
  • the at least one processor may obtain the audio spatial characteristics based on at least one visual cues in the image data, wherein the at least one visual cues may include information related with an object corresponding to a sound source.
  • the at least one processor may identify motion sickness degree of a user based on the video frame information associated with the virtual reality environment, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the motion sickness degree.
  • an method for controlling an electronic apparatus comprising: obtaining a content associated with a virtual reality environment, wherein the content may include audio data and image data, obtaining video frame information associated with the image data, wherein the video frame information may include a video frame speed for displaying at least one video frame included in the image data, obtaining audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtaining enhanced audio data by rendering the audio data based on the audio spatial characteristics.
  • the audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
  • the obtaining the audio spatial characteristics may include obtaining an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtaining the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
  • the method may include outputting the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  • the obtaining enhanced audio data may include identifying a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtaining the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
  • a method of enhanced audio rendering for an immersive virtual reality experience includes receiving a variable video frame information associated with a virtual reality environment, the variable video frame information comprising a speed of one or more variable image frames associated with the virtual reality environment. Further, the method includes generating an audio processing module based on the speed of the one or more variable image frames for preserving one or more spatial audio cues and ensuring that the audio processing module maintains a sense of immersion for a user during time-manipulated variable video frame information being experienced within the virtual reality environment.
  • the method includes determining at least one of localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module, in addition to, rendering the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
  • a system of enhanced audio rendering for an immersive virtual reality experience includes a memory, and at least one processor coupled to the memory.
  • the at least one processor is configured to receive variable video frame information associated with a virtual reality environment, the variable video frame information comprising a speed of one or more variable image frames associated with the virtual reality environment.
  • at least one processor is configured to generate an audio processing module based on the speed of the one or more variable image frames for preserving one or more spatial audio cues and ensuring that the audio processing module maintains a sense of immersion for a user during time-manipulated variable video frame information being experienced within the virtual reality environment.
  • At least one processor is configured to determine at least one of localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module.
  • at least one processor is configured to render the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
  • Figure 1 illustrates a block diagram depicting an embodiment of an audio engine, in accordance with an embodiment of the present disclosure
  • Figure 2 illustrates a block diagram of an environment of enhanced audio rendering for an immersive virtual reality experience for a user within a virtual environment, in accordance with an embodiment of the present disclosure
  • Figure 3 illustrates a block diagram of a system, in accordance with an embodiment of the present disclosure
  • Figure 4 illustrates a flow chart depicting a method of enhanced audio rendering for an immersive virtual reality experience for a user within a virtual environment, in accordance with an embodiment of the present disclosure
  • Figure 5 illustrates a flow chart depicting a method for improving voice quality using a vocoder of the system, according to embodiments disclosed herein;
  • Figure 6A illustrates an example block diagram depicting a noise identification model of a classification module, according to the embodiments disclosed herein;
  • Figure 6B illustrates an example diagram depicting an embodiment of the classification module, according to embodiments disclosed herein;
  • Figure 7A illustrates a block diagram depicting an object-based identification model, according to embodiments disclosed herein;
  • Figure 7B illustrates a block diagram depicting a multimodal analysis of videos of the object-based identification model, according to embodiments disclosed herein;
  • Figure 7C illustrates another block diagram depicting an object-based identification model, according to embodiments disclosed herein;
  • Figure 8 illustrates an example diagram depicting the localization and movement module, according to embodiments disclosed herein;
  • Figure 9 illustrates another example diagram depicting a time-scale modification pipeline of a spatial effect module, according to embodiments disclosed herein;
  • Figure 10A illustrates an example block diagram depicting a resampling module, according to the embodiments disclosed herein.
  • Figure 10B illustrates another example block diagram depicting the resampling module, according to the embodiments disclosed herein.
  • Figure 11 illustrates an example block diagram control method for the electronic apparatus.
  • modules or engines that carry out a described function or functions.
  • modules or engines which may be referred to herein as units or blocks or the like, or may include blocks or units, may be physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software.
  • the circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like.
  • circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block.
  • a processor e.g., one or more programmed microprocessors and associated circuitry
  • Each block may be physically separated into two or more interacting and discrete blocks without departing from the scope of the inventive concepts.
  • the blocks may be physically combined into more complex blocks without departing from the scope of the inventive concepts.
  • FIG. 1 illustrates a block diagram 100 depicting an embodiment of an audio engine, in accordance with an embodiment of the present disclosure.
  • an audio engine 101 may be communicated with a Virtual Reality (VR) engine 106.
  • the audio engine 101 may include an audio rendering engine 114, and an audio processing module 116.
  • the VR engine 106 updates a visual scene to match the new viewpoint.
  • the audio engine 101 may be adapted to update with the same speed and accuracy.
  • the audio engine 101 dynamically may be adapted to adjust audio playback in real-time to match the visual changes detected by the VR engine 106.
  • the audio engine 101 may include advanced spatial audio rendering techniques to accurately simulate the direction and distance of sound sources relative to the user’s position and movements. Further, the audio engine 101 may be adapted to maintain constant communication with the VR engine 106, receiving real-time updates about the visual changes and adjusting the audio output accordingly.
  • VR controllers 102 may track hand movements and provide input to an operating system 104.
  • the operating system 104 may include a device driver(s) 110 and an original original Software Development Kit (SDK)) 112 adapted to manage hardware resources and provide a platform for VR applications.
  • SDK Software Development Kit
  • the device drivers 110 and the original SDK 112 may be configured to facilitate communication between the hardware and software components of the VR engine 106.
  • a VR content 108 may represent digital content and scenes that the user experiences within the virtual environment.
  • FIG. 2 illustrates a block diagram 200 of an environment of enhanced audio rendering for an immersive virtual reality experience for a user 203 within a virtual environment, in accordance with an embodiment of the present disclosure.
  • the user 203 is using a wearable device 202.
  • the wearable device 202 is a head-mounted display device configured to display a VR environment.
  • the wearable device 202 may be configured to generate audio and video data to enable the user 203 to experience the VR environment.
  • the wearable device 202 may be connected to a system 204 configured to provide the enhanced audio rendering for the wearable device 202.
  • the wearable device 202 may be connected to a system 204 configured to provide the enhanced audio rendering for the wearable device. Then explain the system 204 that it may be located remotely or within the wearable device. Further, it should be noted that although Figure 2 depicts the system 204 as separate from the wearable device 202. In one embodiment, the system 204 may be integrated into the wearable device 202.
  • the system 204 may include the VR engine 106, the audio rendering engine 114, the audio processing module 116, a scene change detector 118, a viewport change detector 206, and a point of view prediction 208.
  • the VR engine 106 may be adapted to create and manage the virtual environment in which the user interacts.
  • the audio rendering engine 114 may be adapted to process and generate audio in the VR environment.
  • the audio rendering engine 114 plays a critical role in maintaining audio-visual synchronization and providing an immersive auditory experience.
  • the audio processing module 116 may handle detailed processing and manipulation of audio signals.
  • the scene change detector 118 may be adapted to continuously monitor the VR environment for significant changes in the visual scene.
  • the significant changes may include transitions between different scenes, major shifts in visual content, or changes in the VR environment that require corresponding adjustments in audio rendering.
  • the scene change detector 118 identifies when the user 203 moves from one scene to another or when a significant event occurs within the scene, such as entering a new room or experiencing an explosion within the VR environment.
  • the viewport change detector 206 may be configured to track changes in the user’s viewport within the VR environment.
  • the track changes may include monitoring the user’s head movements and the corresponding changes in the visible portion of the VR scene.
  • the viewport change detector 206 may be configured to inform the audio engine 101 to modify spatial audio rendering based on the viewpoint, ensuring that sounds are accurately localized and move naturally with the user’s head movements.
  • the point of view prediction 208 may be adapted to predict the future head movements of the user 203 and changes in viewpoint.
  • the point of view prediction 208 may help preemptively adjust the audio rendering to maintain synchronization with the visual content.
  • the wearable device 202 may be configured to receive variable video frame information associated with the virtual reality environment.
  • the variable video frame information may include information such as, but not limited to, a speed of variable image frames associated with the virtual reality environment.
  • the system 204 may be configured to may be configured to generate the audio processing module 116 based on the speed of the variable image frames for preserving spatial audio cues and ensuring that the audio processing module 116 maintains the sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment.
  • the audio processing module 116 may include the audio features based on the received variable video frame information.
  • the spatial audio cues provide information about the location and the movement of sound sources in the virtual environment, the spatial audio cues include speed up or speed down at which sound is coming.
  • the system 204 may be configured to determine localization and movement of sounds in the virtual environment based on the speed of the variable image frames, and the generated audio processing module 116.
  • the movement of sounds may be simulated based on the speed of the variable image frames associated with the virtual reality environment and the determined localization and movement of sounds.
  • the system 204 may be configured to determine the movement of sounds based on the speed of variable image frames associated with the virtual reality environment and changes in visual cues, the visual cues within a current scene provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds.
  • the system 204 may be configured to render the enhanced audio in the VR environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
  • render the enhanced audio may include rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects.
  • the rendering enhanced audio may provide synchronization of audio with respect to the speed of the variable image frames associated with the virtual reality environment.
  • system 204 may be configured to identify spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined localization and movement of sounds.
  • the spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203.
  • system 204 may be configured to update the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment.
  • system 204 may be configured to predict motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment. Further, the system 204 may be configured to determine the movement of the user 203 based on the variable video frame information associated with the virtual reality environment.
  • system 204 may include extracting features corresponding to the variable video frame information for noise, speech, music, and object tagging. Furthermore, the system 204 may include a plurality of audio features associated with the speed of the variable image frames.
  • the VR content 108 may represent digital content and scenes that the user 203 experiences within the virtual environment.
  • the VR content 108 may include both visual and audio elements to create an immersive experience.
  • the scene change detector 118 may be responsible for monitoring the current VR scene and determining any changes or transitions. For example, the scene change detector 118 may identify when the visual content changes significantly, such as moving to a new scene or a significant event occurring within the scene. Further, the scene change detector 118 may predict the user’s future head movements and adjust the audio accordingly to maintain synchronization.
  • the system 204 uses data from the scene change detector 118 to assess the potential for motion sickness and make necessary adjustments to the audio rendering to mitigate it.
  • the system 204 detects any changes in the speed of the video content based on the predicted scene changes.
  • the system 204 may monitor and detect changes in video playback speed that may be required to keep the audio and visual elements in synchronization.
  • Sound process the system 204 may be adapted to determine localization and movement of sound based on changes in video speed.
  • Audio rendering the system 204 may provide time-manipulated audio based on user orientation and speed of visual cues.
  • immersive content the system 204 may provide enhanced immersive audio to reduce the motion sickness of the user 203 in the current scene.
  • Figure 3 illustrates a block diagram of system 204, in accordance with an embodiment of the present disclosure.
  • Figure 4 illustrates a flow diagram depicting a method 400 of enhanced audio rendering for an immersive virtual reality experience for the user 203 within a virtual environment, in accordance with an embodiment of the present disclosure.
  • Figures 2, 3, and 4 are explained in conjunction with each other.
  • the system 204 may include, but is not limited to, a processor 304, memory 302, an interface 306, and a plurality of modules 308.
  • the memory 302, the interface 306, and the plurality of modules 308 may be coupled to the processor 304.
  • the processor 304 may be a single processing unit or several units, all of which could include multiple computing units.
  • the processor 304 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and/or any device that manipulates signals based on operational instructions.
  • the processor 304 is configured to fetch and execute computer-readable instructions and data stored in the memory 302.
  • the memory 302 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and/or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. Further, the memory 302 may include an operating system 312 for performing one or more tasks of the system 204, as performed by a generic operating system 312 in the communications domain.
  • the processor 304 may be configured to receive variable video frame information associated with a virtual reality environment.
  • the variable video frame information may include the speed of variable image frames associated with the virtual reality environment.
  • the processor 304 may be configured to generate the audio processing module 116 based on the speed of the variable image frames for preserving the spatial audio cues and ensuring that the audio processing module 116 maintains a sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment.
  • the spatial audio cues may provide information about the location and the movement of sound sources in the virtual environment.
  • the spatial audio cues may include the speed up or the speed down at which sound is coming.
  • the processor 304 may be configured to determine the localization and movement of sounds in the virtual environment based on the speed of the variable image frames, and the generated audio processing module 116. Furthermore, the processor 304 may be configured to render the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds. The movement of sounds may be simulated based on the speed of the variable image frames associated with the virtual reality environment and the determined localization and movement of sounds.
  • the rendering enhanced audio may include rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects, the rendering enhanced audio may provide the synchronization of audio with respect to the speed of the variable image frames associated with the virtual reality environment.
  • the processor 304 may be configured to identify spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined at least one of localization and movement of sounds, wherein the spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203.
  • the processor 304 may be configured to update the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment.
  • the audio processing module 116 may include the plurality of audio features based on the received variable video frame information. The plurality of audio features may be associated with the speed of the variable image frames.
  • the processor 304 may be configured to determine the movement of sounds based on the speed of the variable image frames associated with the virtual reality environment and the changes in visual cues.
  • the visual cues within the current scene may provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds.
  • the processor 304 may be configured to predict the motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment.
  • the processor 304 may be configured to determine the movement of the user 203 based on the variable video frame information associated with the virtual reality environment.
  • the processor 304 may be configured to extract features corresponding to the variable video frame information for noise, speech, music, and object tagging.
  • the plurality of modules 308 amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types.
  • the plurality of modules 308 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and/or any other device or component that manipulates signals based on operational instructions.
  • the plurality of modules 308 may be configured to perform the steps of the present disclosure using the data stored in a database 310 for automated parking management, as discussed herein.
  • the database 310 may be configured to store the information as required by the plurality of modules 308 and the one or more processors 304 for enhanced audio rendering for the immersive virtual reality experience for the user 203 within the virtual environment.
  • the plurality of modules 308 can be implemented in hardware, instructions executed by a processing unit, or by a combination thereof.
  • the processing unit can comprise a computer, a processor, such as the processor 304, a state machine, a logic array, or any other suitable wearable device capable of processing instructions.
  • the processing unit can be a general-purpose processor which executes instructions to cause the general-purpose processor to perform the required tasks or, the processing unit can be dedicated to performing the required functions.
  • the plurality of modules 308 may be machine-readable instructions (software) which, when executed by a processor/processing unit, perform any of the described functionalities.
  • the plurality of modules 308 may include a set of instructions that may be executed to cause the system 204 to perform any one or more of the methods disclosed herein.
  • the plurality of modules 308 may be configured to perform the steps of the present disclosure using the data stored in the memory 302 to provide enhanced audio rendering for the immersive virtual reality experience for the user 203 within the virtual environment, as discussed throughout this disclosure.
  • each of the modules 308 may be hardware units that may be outside the memory 302.
  • the plurality of modules 308 may include a detection module 310, a prediction module 313, the audio rendering engine 114, and the audio processing module 116.
  • the audio rendering engine 114 may include a spatial effect module 314, a resampling module 316, and a vocoder 318.
  • the audio processing module 116 may include sub-modules such as an audio feature extraction module 320, a classification module 322, and a localization and movement module 324.
  • the detection module 310 may be configured to identify changes and events within the VR environment that require adjustments in audio and visual rendering.
  • the detection module 310 may be configured to continuously monitor the VR scene and detect user interactions, head movements, and scene transitions. For example, the detection module 310 monitors and identifies changes in the user’s viewpoint within the VR environment.
  • the detection module 310 continuously tracks the head movements of the user 203 to determine changes in the viewing angle and direction.
  • the detection module 310 may detect when the user’s viewport (for example, the visible area of the virtual environment) shifts due to head movements or other inputs.
  • the detection module 310 may be configured to identify significant changes within the VR environment.
  • the detection module 310 may be configured to detect specific events or changes that indicate a scene transition, such as entering a new room, an explosion, or the appearance of new objects.
  • the prediction module 313 may be configured to anticipate future movements and actions of the user 203 based on current and past data. By analyzing patterns in the user's behavior and movements, the prediction module 313 may be configured to predict the user’s next position or action. The prediction module 313 may be adapted to pre-render images and pre-process audio, reducing latency and ensuring smoother transitions and more accurate synchronization between the audio and visual elements.
  • the spatial effect module 314 may be configured to apply spatial effects to the audio signals, creating a 3D soundscape that matches the virtual environment.
  • the spatial effect module 314 may adjust the direction, distance, and movement of sounds to align with the user’s perspective and the visual scene.
  • the resampling module 316 may be configured to adjust the sampling rate of audio signals to ensure they match the current playback speed and synchronization requirements.
  • the vocoder 318 may process and manipulate the audio signals to alter pitch and timing to improve the quality.
  • the vocoder 318 ensures that voice and other audio effects remain natural and intelligible, even when there are changes in playback speed.
  • the Vocoder 318 may be a type of vocoder-purposed algorithm which is used to interpolate information present in the frequency and time domains of audio signals by using phase information extracted from a frequency transform.
  • the audio feature extraction module 320 may be configured to extract acoustic features from audio signals (for example, a speech signal), such as pitch, tone, and amplitude.
  • the features may be used for further processing and analysis thereby ensuring that the audio remains high-quality and immersive.
  • the acoustic features may include Linear Predictive Cepstral Coefficients (LPCC) features and Mel-Frequency Cepstral Coefficients (MFCC) features.
  • LPCC features may include, but are not limited to, 13 delta LPCC features, 13 delta-delta LPCC features, 13 LPCC features, and the like.
  • the MFCC features may include 12 MFCC Cepstral Coefficients, 12 Delta MFCC features, 12 Double Delta MFCC features, 1 energy coefficient, 1 delta energy coefficient, 1 double Delta energy coefficient, and the like.
  • Combining LPCC and MFCC features may provide a more comprehensive representation of the speech signal.
  • Principal Component Analysis (PCA) By applying Principal Component Analysis (PCA) to the combined features, dimensionality may be reduced while preserving the important information.
  • PCA Principal Component Analysis
  • the combined use of LPCC and MFCC features, along with the PCA, may enhance the capability of the system 204 to analyze and process speech signals within the VR environment.
  • the classification module 322 may be configured to categorize the audio signals based on features and context within the VR environment.
  • the classification module 322 may be configured to identify different types of sounds such as dialogue, ambient noise, and sound effects to apply processing techniques.
  • the localization and movement module 324 may be configured to determine spatial positions and movements of sound sources within the virtual environment. The localization and movement module 324 may ensure that audio cues are accurately positioned and move in harmony with the visual scene, enhancing the sense of presence and immersion for the user 203.
  • the modules 310, 312, 114, and 116 may be in communication with each other.
  • the modules 310, 312, 114, and 116 may be a part of the processor 304.
  • the processor 304 may be configured to perform the functions of the modules 310, 312, 114, 116.
  • At least one of the modules 310, 312, 114, and 116 may be implemented through an artificial intelligence (AI) model.
  • AI artificial intelligence
  • a function associated with AI model may be performed through the non-volatile memory, the volatile memory, and the processor 304.
  • the processor 304 may include one or a plurality of processors.
  • one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and/or an AI-dedicated processor such as a neural processing unit (NPU).
  • the one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory.
  • the predefined operating rule or artificial intelligence model is provided through training or learning.
  • the learning may be performed in a device in which AI model according to an embodiment is performed, and/or may be implemented through a separate server/system.
  • the AI model may include a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights.
  • Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
  • the learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction.
  • Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
  • system 204 may be a part of the wearable device 202.
  • the system 204 may be connected to the wearable device 202.
  • the wearable device 202 may be a virtual reality (VR) headset designed to provide the immersive VR experience.
  • the wearable device 202 may include integrated audio and visual systems that allow the user 203 to interact with and perceive the virtual environment.
  • Figure 4 illustrates the flow chart depicting a method 400 of the enhanced audio rendering for an immersive virtual reality experience for the user within a virtual environment, according to embodiments disclosed herein.
  • the method 400 may be performed at the wearable device 202.
  • the method 400 may be performed by the system 204 comprising the processor 304 and the memory 302.
  • the method 400 includes receiving, at the wearable device 202, variable video frame information associated with the virtual reality environment, the variable video frame information may include information such as, but not limited to, a speed of variable image frames associated with the virtual reality environment.
  • the method 400 includes generating the audio processing module 116 based on the speed of the variable image frames for preserving spatial audio cues and ensuring that the audio processing module 116 maintains the sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment.
  • the audio processing module 116 includes the audio features based on the received variable video frame information.
  • the spatial audio cues provide information about the location and the movement of sound sources in the virtual environment, the spatial audio cues include speed up or speed down at which sound is coming.
  • the method 400 includes determining the localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module 116.
  • the movement of sounds is simulated based on the speed of the one or more variable image frames associated with the virtual reality environment and the determined localization and movement of sounds. Further, determining the movement of sounds based on the speed of variable image frames associated with the virtual reality environment and changes in visual cues, the visual cues within a current scene provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds.
  • the method 400 includes rendering the enhanced audio in the VR environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
  • the method 400 includes identifying spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined localization and movement of sounds.
  • the spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203.
  • the method includes updating the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment.
  • the method 400 includes predicting motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment. Further, the method 400 includes determining the movement of the user 203 based on the variable video frame information associated with the virtual reality environment.
  • the method 400 includes rendering the enhanced audio includes rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects.
  • the rendering enhanced audio provides synchronization of audio with respect to the speed of the one or more variable image frames associated with the virtual reality environment.
  • the method 400 includes extracting features corresponding to the variable video frame information for noise, speech, music, and object tagging.
  • the method 400 includes the plurality of audio features associated with the speed of the variable image frames.
  • the audio feature extraction module 320 may be configured to process the audio signal to extract the features for further processing.
  • the steps involve several transformations and computations, which are detailed below.
  • Figure 5 illustrates the flow chart depicting a method 500 for improving voice quality using the vocoder of the system 204, according to embodiments disclosed herein.
  • the method 500 may be performed by the system 204 comprising the processor 304 and the memory 302 of the system 204.
  • the method 500 includes accepting the raw audio waveform (speech) as the initial input.
  • the method 500 includes transforming the audio waveform into a mel-spectrogram representation to capture both frequency and time information.
  • the method 500 includes normalizing the Mel-spectrogram to ensure consistent data for further processing.
  • the method 500 includes converting the normalized Mel-spectrogram into the format required by the WaveNet decoder.
  • the method 500 includes generating audio samples one by one in a casual autoregressive manner, using previously generated samples and mel-spectrogram information to predict the next sample.
  • the method 500 includes applying an inverse filter to the generated audio to remove any pre-emphasis applied during preprocessing, ensuring a more natural sound.
  • the method 500 includes adjusting the audio signal to an appropriate volume level to ensure consistency and enhance the listening experience.
  • the method 500 includes delivering the improved speech audio waveform with enhanced quality.
  • the various operations of methods described above may be performed by any suitable device capable of performing the operations, such as the processing circuitry discussed above.
  • the operations of methods described above may be performed by various hardware and/or software implemented in some form of hardware (e.g., processor, ASIC, etc.).
  • FIG. 6A illustrates an example block diagram 600 depicting a noise identification model of the classification module 322, according to the embodiments disclosed herein.
  • the block diagram 600 includes a Residual Neural Network (ResNet) module 602, a Visual Geometry Group VGGish (for example, CNN Algorithm) module 604, a Neural Network (NN) classification module 606, and noise segments module 608.
  • Residual Neural Network Residual Neural Network
  • VGGish for example, CNN Algorithm
  • NN Neural Network
  • the input 610a may be a single video
  • the output may be the noise segments 608.
  • the noise segments 608 may include an identification of noise or not or corresponding sound.
  • the ResNet module 602 may be adapted to process video frames to extract deep image features.
  • the deep image features may capture the spatial and temporal context of the video.
  • the VGGish module 604 may be required for audio classification tasks.
  • the VGGish module 1004 may be adapted to convert the audio features into a format suitable for neural network processing.
  • the NN classification module 606 may be configured to take combined features from the ResNet module 602 and the VGGish module 604.
  • the NN classification module 606 may be configured to perform the final classification to determine if the segment is “noise” or “not noise”.
  • the noise segments 608 may be the final output consists of the classification results for each segment. Each segment may include noise and not noise. The noise may indicate the presence of noise and not noise may indicate an absence of noise corresponding to sound.
  • FIG. 6B illustrates an example diagram depicting an embodiment of the classification module 322, according to embodiments disclosed herein.
  • the classification module 322 may include the audio signal 610b, a music classifier 612, a song extractor 614, and speech, and music sequences 616.
  • the audio signal 610b, the music classifier 612, the song extractor 614, and the speech, and music sequences 616 are also part of the modules 208 of the system 204.
  • the audio signal 610b may be extracted from the video 610a, which serves as the raw data for the classification.
  • the music classifier 612 may include a classification model, specifically a Support Vector Machine (SVM) classifier, which processes the audio signal 610b to distinguish between music and non-music segments.
  • the music and the non-music segments of video m may include segmenting video m into frames with length l, calculating audio features for each frame, classifying video frames in music and non-music classes using a Support Vector Machine, and returning music and non-music segments.
  • the song extractor 614 may be a component that processes segments identified as music to extract song sequences, leveraging the fact that songs are typically accompanied by music.
  • the song extractor 614 may be configured to take music component provided by the music classifier 612 as an input and searches for song sequences in it.
  • the music part certainly includes the songs, but there may also be dialogue and action scenes with background music in the music component.
  • the song extractor 614 may be divided into three sub-components such as a potential song sequence generator 614a, a vocal detector 614b, and a vocal pattern checker 614c.
  • the potential song sequence generator 614a may be configured to accept the music part of a video and generate a number of potential song sequences (PSS), which are highly likely to be songs.
  • PSS potential song sequences
  • the potential song sequence generator 614a may break the music part of the video 610a into multiple sequences where a sequence refers to consecutive musical seconds without break. The potential song sequence generator 614a then applies rules of sequences in order to extract a potential song sequence.
  • the vocal detector 614b may be configured to detect vocal and non-vocal segments of the possible song sequence.
  • the vocal detector 614b may be configured to use the features for music or non-music classification spectral centroid and mel-frequency cepstral coefficients features. Similar to that of the music classifier 612 along with some additional audio features, spectral centroid and Mel-frequency cepstral coefficients may be used.
  • the spectral centroid may be defined as a weighted mean of signal frequencies.
  • the spectral centroid may be determined by a Fourier transform using frequency magnitudes as weights.
  • the vocal pattern checker 614c may be configured to check the pattern of vocal and non-vocal segments of possible song sequences detected by the vocal detector 614b.
  • the vocal pattern checker 614c may use the pattern of vocal and non-vocal segments to differentiate between Potential Song Sequences (PSS) that are songs and the PSS that are action or dialogues.
  • PSS Potential Song Sequences
  • the vocal segment may be either intro, chorus, verse, or outro.
  • the duration for vocal segments may fall in a range between 20 to 100 seconds.
  • the non-vocal segment may be a bridge. Depending on the place of its appearance, the duration of the bridge in non-vocal segments ranges from 20 to 40 seconds.
  • the pattern of vocal and non-vocal segments may be used to differentiate between potential song sequences which are songs to the PSS which are action or dialogues.
  • the performance of the classification of music from audio depends mainly on the selection of appropriate audio features and classifiers.
  • the audio features may include a zero-crossing rate, an audio spectrum flux, a short-time energy or sub-band energy distribution, and an audio intensity.
  • Zero Crossing Rate may be defined as measures the rate at which the audio signal changes sign, indicating the signal’s noisiness or the presence of high frequencies.
  • the audio spectrum flux may be defined ascaptures the rate of change in the power spectrum of the audio signal, indicating the presence of dynamic elements such as music or speech.
  • the short-time Energy or sub-band energy Distribution may be defined as analyzes the energy content within short windows of the audio signal, identifying variations characteristic of music or speech.
  • the audio intensity measures the perceived loudness of the audio signal, helping to distinguish between music and quieter non-music segments.
  • the output of the song extractor 614 may categorize the input audio signal into distinct speech and music sequences 616 for further analysis or processing.
  • the speech and music sequences 616 may provide a clear categorization of the audio signal into speech and music sequences.
  • the speech and music sequences 616 may be used for various applications such as content indexing, music retrieval, and audio analysis.
  • Figure 7A illustrates a block diagram 700a depicting an object-based identification model, according to embodiments disclosed herein.
  • the object-based identification model 702 may be configured to identify structural and semantic relationships between a non-subject or subject audio and the video frames by object identification segments (for example, frame objects tagger) 704.
  • the object-based identification model 702 may include a segmentation module 702a adapted to extract the objects for tagging.
  • the object-based identification model 702 may include a Motion Scale Invariant feature Transform (Mo SIFT) 702b, a Mel-Frequency Cepstral co-efficient 702c, a Bag of Words (BoW) 702d, a SVM classification (for example, Dynamic Time Wrapping (DTW)) 702e, and another SVM classification 702f.
  • Mo SIFT Motion Scale Invariant feature Transform
  • BoW Bag of Words
  • SVM classification for example, Dynamic Time Wrapping (DTW)
  • DTW Dynamic Time Wrapping
  • the object-based identification model 702 may train three two-class SVMs in order to learn tagging models.
  • the first SVM model 702e may be constructed using low-level audio features, and the SVM classification 702f using low-level visual features. Fusing the predictions of the two SVM models i.e., a visual classifier and an audio classifier using another two-class SVM for the final prediction of object identification segments for Frame objects Tagger 704.
  • the Mo SIFT 702b may be a standard SIFT algorithm applied to find visually distinctive interest points in the spatial domain.
  • the Mo SIFT 702b (for example, 256 dimensions) may be designed to represent the feature point in two parts: the first 128 dimensions are standard SIFT features and the remaining 128 dimensions are the aggregated histogram of optical flow.
  • the vocabulary of visual words may be typically defined as the cluster centers obtained from k-means clustering over a large collection of sample low-level descriptors (Mo SIFT descriptors).
  • the BoW 702d represents each video sequence as a histogram over a set of visual words (motion and color key points) to generate a fixed-dimensional encoding that may be processed using the standard classifier.
  • Figure 7B illustrates a block diagram 700b depicting a multimodal analysis of videos of the object-based identification model, according to embodiments disclosed herein.
  • the object-based identification model 700b includes a visual feature extraction 706a, an audio feature extraction 706b, the visual classifier 706c, the audio classifier 706d, and a multimodal classifier 706e.
  • the visual feature extraction 706a may be the process of analyzing visual data from the video to identify key characteristics and patterns.
  • the characteristics and patterns may include processing frames to extract information such as color, texture, shape, and motion.
  • the audio feature extraction 706b may include analyzing the audio signal to identify important characteristics such as pitch, rhythm, timbre, and zero crossing rate.
  • the audio feature extraction 706b may convert raw audio data into a set of meaningful features that may be used for further analysis and classification.
  • the audio feature extraction 706b may include Mel Frequency Cepstral Coeficients (MFCC) that divide each utterance into overlapping frames, each frame is 25ms long with a step of 10ms.
  • MFCC Mel Frequency Cepstral Coeficients
  • the visual classifier 706c may be an algorithm or model that uses extracted visual features to categorize video content into different classes.
  • the visual classifier 706c may use machine learning techniques like convolutional neural networks (CNNs) to recognize and classify objects, scenes, or actions within the video.
  • CNNs convolutional neural networks
  • histograms obtained from BoW are high-dimensional vectors that may be classified based on different scenes using a standard classifier, typically the SVM model.
  • the audio classifier 706d may use extracted audio features to categorize the audio content into different classes.
  • the audio classifier 706d may involve identifying whether a segment contains music, speech, or other types of sounds.
  • the audio classifier 706d may include Dynamic Time Warping (DTW) that matches two signals by calculating the best path that is the best alignment between the two signals resulting in the level of similarity between them.
  • DTW Dynamic Time Warping
  • the SVM may be used to classify the sound using DTW as its distance metrics.
  • the multimodal classifier 706e may be adapted to combine both visual and audio features to make more accurate and robust classifications. By integrating data from both modalities, the multimodal classifier 706e may understand the context and content of the video, leading to improved performance in tasks such as identifying scenes, actions, or specific events. In an embodiment, both the outputs of visual classifier 706c and the audio classifier 706d may be fed to the multimodal classifier 706e. The final output from the multimodal classifier 706e may be whether the video clip has classification and objects tagging.
  • Figure 7C illustrates another block diagram 700c depicting an object-based identification model, according to embodiments disclosed herein.
  • the object-based identification model 700c may use existing networks as building blocks which act as a strong core of model and developing a panoptic head that combines outputs of semantic and instance segmentations.
  • the object-based identification model 700c may include input frame 708a, backbone network 708b, semantic head 708c, Instance Head 708d, a Panoptic head 708e, panoptic logits 708f, object tagging 708g, semantic logits 710a, stuff logits 710b, the panoptic logits 708f.
  • the input frame 708a may be a single image or frame from the video that is fed into the object-based identification model for analysis.
  • the backbone network 708b may be a pre-trained neural network that acts as the core feature extractor.
  • the backbone network 708b may process the input frame 708a to produce high-level feature maps.
  • the backbone networks 708b may include ResNet, VGG, and EfficientNet.
  • the semantic Head 708c may process the feature maps from the backbone network 708b to perform semantic segmentation, which classifies each pixel in the frame into different classes (e.g., road, sky, building).
  • the instance head 708d may process the feature maps to perform instance segmentation, which identifies and delineates individual objects within the frame (e.g., a specific car, a particular person).
  • the panoptic head 708e may combine the outputs from the semantic head 708c and the instance head 708d to create a unified segmentation that includes both “stuff” (background regions like sky, road) and “things” (discrete objects like cars, people). The results in panoptic logits 708f.
  • the panoptic logits 708f may be the output probabilities from the panoptic head 708e, indicating the likelihood of each pixel belonging to a specific class or instance.
  • the object tagging 708f may involve assigning labels or tags to the identified objects based on the panoptic logits.
  • the object tagging 708f may annotate the objects with respective categories (e.g., car, person).
  • the semantic logits 710a may be the output probabilities from the semantic head 708c, indicating the likelihood of each pixel belonging to a specific semantic class.
  • the Stuff logits 710b may be the specific probabilities from the semantic logits that correspond to background regions or “stuff” classes (e.g., sky, grass).
  • instance logits 710c may be the output probabilities from the instance head 708d, indicating the likelihood of each pixel belonging to a specific instance of an object.
  • the output probabilities may include instance i, and instance j.
  • the occlusion head 710d may identify and process occlusions within the frame, where one object may partially or fully block another. This occlusion head 710d may help in accurately segmenting overlapping objects.
  • the occlusion matrix 710f may be a representation of occlusions detected by the occlusion head 710d, showing which parts of objects are occluded and how they overlap with each other.
  • the fusion 710f may refer to the process of combining information from various logits (for example, semantic, instance, occlusion) to produce a coherent and accurate final output.
  • the ground truth mask 710g may be the actual annotated data used for training and validating the model.
  • the edge map head 710h may detect and highlight the edges of objects within the frame, helping to delineate boundaries more precisely.
  • FIG. 8 illustrates an example diagram 800 depicting the localization and movement module, according to embodiments disclosed herein.
  • the localization and movement module 324 may include building audio datasets 802, building interaural phase difference (IPD) datasets 804, and a training model 806.
  • the building audio datasets 802 may include a multichannel impulse response database 802a, a real Room Impulse Response (RIR) 802b, and a simulated Room Impulse Response (RIR) 802c.
  • the building (IPD) interaural phase difference datasets 804 may include short-time Fourier transform (STFT) 804a, IPD feature extraction 804b, real IPD 804c, and a simulated IPD 804d.
  • STFT short-time Fourier transform
  • the training model 806 includes a convolutional neural network plus regression model 806a.
  • Pyroomacoustics platform and the multichannel impulse response database (MIRD) may be used to generate both simulated and real room impulse response (RIR) datasets 804c, 804d.
  • MIRD multichannel impulse response database
  • the multichannel impulse response database 802a may be a collection of impulse responses recorded from multiple channels, providing a comprehensive set of data for different environments and setups.
  • the real RIR 802b may be recorded in real physical rooms, capturing the acoustics and reflections of actual environments.
  • the simulated RIR 802c may be generated through simulations, allowing for the creation of data that may not be easily captured in real-world settings.
  • the STFT 804a may be a mathematical technique applied to audio signals to convert them from the time domain to the frequency domain, facilitating the analysis of frequency components over short time windows.
  • the IPD feature extraction 804b may be the process of extracting interaural phase difference features, which are critical for determining the spatial characteristics of sounds based on the phase differences between signals received at the left and right ears.
  • the real IPD 804c may be derived from real-world recordings, providing authentic spatial audio features.
  • the simulated IPD 804d may be generated through simulations, and used to augment the dataset with controlled variations and scenarios.
  • the IPD features of the sound signal may be firstly extracted from time-frequency domain by the STFT 804a. Then, the IPD features map may be fed to the CNN-R model 806 as an image for sound source localization.
  • the machine learning model 806a combines convolutional neural networks (CNNs) with regression techniques.
  • the CNN part may be used for extracting spatial features from the audio data, while the regression part predicts the precise localization and movement parameters of the sound sources.
  • the CNN with a regression model (CNN-R) to estimate the sound source angle and distance based on the acoustic characteristics of the interaural phase difference (IPD).
  • Figure 9 illustrates another example diagram 900 depicting a time-scale modification pipeline of the spatial effect module 314, according to embodiments disclosed herein.
  • the example diagram 900 may include detailed steps and processes involved in modifying the time scale of an audio signal to maintain its frequency content and ensure high-quality audio output.
  • the example diagram 900 may include the input audio 610b, STN Decomposition 906, a phase vocoder 908, Stretch Sines and Transients, Extract CQT (Constant-Q Transform) Spectrogram, Train WaveNet 912, and envelop detection 910a, 910b, and time stretched output signal 912.
  • the input audio signal may be decomposed as the STN, sines (tonal content), transients (impulsive events), and noise (sound nuances).
  • the sines represent the tonal content of the audio, and transients represent impulsive events or sudden changes in the audio, the noise Represents the background sound nuances.
  • the sines may be processed via a phase vocoder 908 with an identity phase locking.
  • the phase vocoder 908 may be a type of vocoder-purposed algorithm which is used to interpolate information present in the frequency and time domains of audio signals by using phase information extracted from a frequency transform.
  • the vocoder-purposed algorithm may allow frequency-domain modifications to a digital sound file.
  • the noise component may be neutrally resynthesized at the desired TSM factor. Transients may be preserved and relocated onto the new time axis according to detected peaks, which are isolated due to the nature of the STN decomposition.
  • the envelope detection 910 may be configured to reshape the STN according to the envelope of the original signal to compensate for the pre-echo effect.
  • the neural network (WaveNet) 912 may be trained using a CQT spectrogram and the desired TSM factor. The training may enable the network to generate new noise components that match the modified time scale. Using the trained WaveNet, new noise components may be synthesized to align with the stretched or compressed audio signal, maintaining the original spectral qualities. The processed sines, transients, and the synthesized noise may be combined to form a cohesive audio signal.
  • a final processed audio signal 914 may be delivered, which has been time-stretched or compressed while preserving its frequency content and maintaining high audio quality.
  • FIG 10A illustrates an example block diagram 1000 depicting a resampling module 316, according to the embodiments disclosed herein.
  • the example block diagram 1000 may include three blocks 1002, 1004, and 1006.
  • the audio model may represent the original audio signal and its properties before any adjustments are made.
  • the audio model may include the raw audio data that needs to be adjusted in response to changes in the video speed.
  • the resampling module 316 may be configured for adaptive sample rate which aims to adjust the audio playback speed to match the modified video speed while preserving pitch and timbre.
  • the resampling module 316 may include the output audio signals that have been resampled to match the new video speed.
  • the output audio signals may be the final product of the resampling module 316, ready to be synchronized with the modified video playback.
  • Figure 10B illustrates another example block diagram depicting the resampling module, according to the embodiments disclosed herein.
  • audio sampling is the process of transforming a musical source into a digital file. The more samples you take - known as the ‘sample rate’ - the more closely the final digital file may resemble the original. A higher sample rate tends to deliver a better-quality audio reproduction.
  • Sample rates may be usually measured per second, using kilohertz (kHz) or cycles per second.
  • the CDs are usually recorded at 44.1kHz - which means that every second, 44,100 samples were taken. Common sample rates are 44.1khz, 48khz, 88.2khz, 96khz and so on.
  • the resampling module 316 starts by checking a video frame rate 1008. For frame rates 1008 less than 120 frames per second (fps), the audio sample rate is set to 48 kHz 1010. Further, for frame rates higher than 120 fps, pitch changes 1012 may be considered to maintain synchronization. If pitch changes are instantaneous, the resampling module 316 adjusts accordingly. For non-instantaneous changes, the audio latency 1016 may be considered. Medium latency scenarios, for example, speech utterances 1018 use a sample rate of 96 kHz, while high latency scenarios i.e., number of audio sources 1020 use 192 kHz to preserve audio quality.
  • Background noise levels i.e., ambient environment noise 1014 may be considered to ensure that the chosen sample rate provides clear and accurate audio reproduction.
  • the resampling module 316 adapts the sample rate based on the user’s scenario and the audio artifacts affected. Further, the resampling module 316 may include a decision tree-based approach implemented to uniquely identify and adjust each user where multiple rule-based classifiers are used to learn (dataset). The outputs of the decision tree may be then used to predict the correct output.
  • a training phase of the vocoder includes the speech signal, the feature extraction block, the time resolution adjustment block, and a mapping technique block.
  • the testing phase may include the speech signal, the feature extraction block, a wavenet conversion model, a time-invariant synthesis filter, and a synthesized speech.
  • the speech signal may be the initial input in the form of raw audio waveform (for example, speech).
  • the feature extraction block convert the audio waveform into a mel-spectrogram to capture frequency and time information.
  • the time resolution adjustment block adjust the time resolution of the extracted features.
  • Mapping Technique Block apply mapping techniques to correlate the features with the desired audio output characteristics.
  • the speech signal may be initial input for testing, similar to the training phase.
  • the vocoder may transform the audio waveform into the mel-spectrogram for consistency with the training phase.
  • At wavenet conversion model uses the WaveNet architecture to generate audio samples based on the mel-spectrogram.
  • At time-invariant synthesis filter apply a synthesis filter to ensure that the generated audio maintains a natural and consistent sound.
  • synthesized speech the final output, which is the enhanced and improved speech audio waveform.
  • Figure 11 illustrates an example block diagram control method for the electronic apparatus.
  • an electronic apparatus may include a memory, at least one processor comprising a processing circuit, wherein the at least one processor configured to obtain a content associated with a virtual reality environment (1110), wherein the content includes audio data and image data, obtain video frame information associated with the image data (1120), wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data, obtain audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed (1130), and obtain (or render) enhanced audio data by rendering the audio data based on the audio spatial characteristics (1140).
  • the content may refer to data associated with a virtual reality environment, such as audio data, image data, and metadata that collectively define a virtual experience.
  • the content may include video frame information associated with the image data.
  • the at least one processor may obtain the video frame information included in the content.
  • the video frame information may refer to data associated with one or more video frames, including a frame speed, timestamps, frame order, or other temporal parameters used for processing or rendering the video content.
  • the video frame information may be corresponding to the variable video frame information.
  • the video frame information may be represented as variable video frame information video frame metadata, frame timing information, video frame parameters, image frame attributes, or frame display information.
  • the video frame speed may refer to the display rate of video frames over time.
  • the video frame speed may corresponding to the speed of one or more variable image frames.
  • the video frame speed may be represented as a frame display rate, video playback speed, frame rate, visual playback rate, or image sequence speed.
  • the audio spatial characteristics may refer to properties of audio signals that define the perceived location, movement, and directionality of the sound source within a three-dimensional space.
  • the audio spatial characteristics may be represented as spatial sound properties, 3D audio attributes, sound spatial features or audio parameters.
  • the audio spatial characteristics may be corresponding to the at least one of localization and movement of sounds.
  • the audio spatial characteristics may be used for preserving one or more spatial audio cues.
  • the audio spatial characteristics may be used for ensuring that the audio processing model maintains a sense of immersion for a user during time-manipulated video frame information being experienced within the virtual reality environment.
  • the enhanced audio data may refer to audio data that has been processed or rendered to reflect spatial characteristics, improve clarity, or increase immersion within a virtual environment.
  • the enhanced audio data may refer to audio output data generated by applying spatial rendering or sound optimization techniques based on the virtual reality context.
  • the enhanced audio data may be represented as processed audio output, rendered sound data, spatialized audio stream, augmented audio content, or optimized audio signal.
  • the audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
  • the sound localization may refer to the process or ability to determine the position or direction of the sound source within a virtual space, based on auditory cues.
  • the sound localization may be represented as audio source positioning, sound source detection, acoustic localization, directional sound identification, or spatial audio pinpointing.
  • the sound movement may refer to the perceived change in position of the sound source within a virtual space.
  • the sound movement may refer to dynamic changes in the location or direction of the sound source.
  • the sound movement may be represented as audio source motion, dynamic sound positioning, moving sound effect, spatial audio transition, or sound trajectory.
  • the at least one processor may obtain an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtain the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
  • the audio processing model may be corresponding to the audio processing module (116).
  • the audio processing model may obtain a plurality of audio features based on the video frame information.
  • the at least one processor may obtain the audio spatial characteristics indicating the plurality of audio features based on the audio data and the video frame information.
  • the at least one processor may generate the audio spatial characteristics by using the audio processing model.
  • the plurality of audio features may be associated with the video frame speed.
  • the at least one processor may output the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  • the at least one processor may generate the enhanced audio data, which has been processed based on spatial audio characteristics (e.g., direction, distance, movement of sound sources).
  • spatial audio characteristics e.g., direction, distance, movement of sound sources.
  • At least one processor may output enhanced audio data to a speaker of the electronic apparatus based on synchronization information, while image data may be displayed on a display of the electronic apparatus.
  • the enhanced audio data may reflect spatial characteristics of sound, such as direction, distance, and movement, and the image data may be associated with a virtual reality environment.
  • the at least one processor may also receive synchronization information, such as timestamps or frame rate data, that indicates how the audio and image data are temporally related.
  • synchronization information such as timestamps or frame rate data
  • the at least one processor may control output timing of the enhanced audio data to align with the image data. For example, when a sound source appears in a specific frame, the corresponding sound may be output at the exact time the frame is displayed.
  • the enhanced audio data may provide the realism and immersion of the user experience in a virtual environment.
  • the at least one processor may identify a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtain the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
  • the sound spatial position may refer to a location of a sound source in three-dimensional space within a virtual reality environment, identified based on the video frame information and the audio spatial characteristics.
  • the sound spatial position is a specific data output representing the physical or virtual coordinates of a sound source.
  • the Audio spatial characteristics influence perception, while the sound spatial position provides a basis for rendering accurate sound placement.
  • the audio spatial characteristics may represent a set of extracted or derived parameters from audio data, such as directionality vectors, interaural time differences (ITD), interaural level differences (ILD), or head-related transfer function (HRTF) profiles. These parameters are used to model how sound should be rendered to simulate a spatial environment.
  • ITD interaural time differences
  • ILD interaural level differences
  • HRTF head-related transfer function
  • the sound spatial position may refer to a concrete computation result (typically a set of three-dimensional coordinates (x, y, z)) that defines the specific location of a sound source in the virtual environment.
  • the at least one processor may obtain the sound spatial position based on both the audio spatial characteristics and additional contextual data such as video frame information or user orientation.
  • the at least one processor may use audio spatial characteristics as input features for spatial modeling.
  • the at least one processor may use the sound spatial position for rendering spatial audio aligned with visual content.
  • the sound spatial position may indicate a location of the sound source in three-dimensional space relative to a user's position.
  • the at least one processor may update the sound spatial position in real-time based on a change of the video frame speed.
  • the at least one processor may identify the change of the video frame speed. Based on the change of the video frame speed, the at least one processor may update the sound spatial position based on the changed video frame speed.
  • the at least one processor may obtain the audio spatial characteristics based on one or more spatial audio cues in the audio data, wherein the one or more spatial audio cues may include at least one of increasing or decreasing speed associated with a sound source.
  • the at least one processor may obtain the audio spatial characteristics based on at least one visual cues in the image data, wherein the at least one visual cues may include information related with an object corresponding to a sound source.
  • the at least one visual cues within a current scene may provide information about the virtual reality environment and help the user make sense of spatial relationships, motion, and positioning of objects and sounds.
  • the at least one processor may identify motion sickness degree of a user based on the video frame information associated with the virtual reality environment, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the motion sickness degree.
  • the at least one processor may identify(or predict) the motion sickness of the user based on the image data.
  • the at least one processor may analyze frames of the image data to predict a user's degree of motion sickness based on visual motion patterns or frame characteristics.
  • the electronic apparatus may store data in advance to determine whether the pattern is likely to cause motion sickness.
  • the electronic apparatus may analyze a pattern representing screen transitions based on image frames and the stored data.
  • the at least one processor may obtain movement degree of a user.
  • the at least one processor may obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the movement degree.
  • the at least one processor may obtain a sensing data for user’s motion.
  • the electronic apparatus may include a sensor configured to sense a user’s motion.
  • the sensor is IMU (Inertial Measurement Unit), accelerometer sensor, gyroscope sensor or image sensor.
  • the at least one processor may obtain the movement degree of the user based on the sensing data.
  • an method for controlling an electronic apparatus comprising: obtaining a content associated with a virtual reality environment, wherein the content may include audio data and image data, obtaining video frame information associated with the image data, wherein the video frame information may include a video frame speed for displaying at least one video frame included in the image data, obtaining audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtaining enhanced audio data by rendering the audio data based on the audio spatial characteristics.
  • the audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
  • the obtaining the audio spatial characteristics may include obtaining an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtaining the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
  • the method may include outputting the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  • the obtaining enhanced audio data may include identifying a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtaining the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
  • the software may comprise an ordered listing of executable instructions for implementing logical functions, and may be embodied in any "processor-readable medium" for use by or in connection with an instruction execution system, apparatus, or device, such as a single or multiple-core processor or processor-containing system.
  • a software module may reside in Random Access Memory (RAM), flash memory, Read Only Memory (ROM), Electrically Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD ROM, or any other form of storage medium known in the art.
  • RAM Random Access Memory
  • ROM Read Only Memory
  • EPROM Electrically Programmable ROM
  • EEPROM Electrically Erasable Programmable ROM
  • registers hard disk, a removable disk, a CD ROM, or any other form of storage medium known in the art.
  • Embodiments of the present disclosure provide a system that minimizes distortion and unnatural sounds often associated with time-warped video playback.
  • the system preserves the immersive and believable spatial audio experience crucial for VR or 360° video. Further, the system reduces motion sickness by eliminating audio-visual discrepancies that can contribute to discomfort.
  • a time-manipulated variable video is played to reduce motion sickness and the audio is not distorted. While sickness is managed through a low-latency image rendering to the latest viewport, appropriate audio is not generated in this manner to preserve the immersive experience of the AR or VR.
  • the system tracks sound localization and movement in response to video speed changes. Further, the system adapts audio playback speed, resamples audio signals, and applies spatial audio effects for seamless immersion. Furthermore, the system addresses the challenge of maintaining immersive audio during time-manipulated VR and mitigates motion sickness.
  • Embodiments disclosed herein may be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements.
  • the elements may be at least one of a hardware device, or a combination of hardware device and software module.
  • Embodiments disclosed herein may be implemented using at least one hardware device and performing network management functions to control the elements.

Landscapes

  • Engineering & Computer Science (AREA)
  • Multimedia (AREA)
  • Physics & Mathematics (AREA)
  • Acoustics & Sound (AREA)
  • Signal Processing (AREA)
  • Stereophonic System (AREA)

Abstract

Disclosed is a method of an enhanced audio rendering for an immersive virtual reality experience. The method includes receiving a variable video frame information associated with a virtual reality environment, the variable video frame information comprising a speed of variable image frames associated with the virtual reality environment. Further, the method includes generating an audio processing module (116) based on the speed of the variable image frames for preserving spatial audio cues and ensuring that the audio processing module maintains a sense of immersion for a user (203) during time-manipulated variable video frame information being experienced within the virtual reality environment, determining localization and movement of sounds in the virtual environment based on the speed of the variable image frames, and the generated audio processing module, and rendering the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds.

Description

SYSTEM AND METHOD OF ENHANCED AUDIO RENDERING FOR AN IMMERSIVE VIRTUAL REALITY EXPERIENCE
The present disclosure generally relates to the field of virtual reality systems, and more specifically relates to a system and method of an enhanced audio rendering for an immersive virtual reality experience to a user within a virtual environment.
Motion sickness in virtual reality (VR) environments is a common issue that many individuals experience. This phenomenon occurs when there is a disconnect between what the eyes see and what the inner ear senses, leading to feelings of nausea, dizziness, and discomfort. When users are exposed to rapid movements or visuals that do not align with their physical movements, the brain can become confused, resulting in the motion sickness. Various techniques have been developed to mitigate the symptoms, primarily such solutions are based on reducing latency, adjusting a field of view, and creating smoother transitions between movements. Low-latency image rendering to the current viewport is a standard solution for managing visual aspects of motion sickness. However, such conventional methods often fail to consider the auditory component, which is crucial for maintaining an immersive AR/VR experience.
In time-manipulated VR and 360° video experiences, current audio rendering techniques struggle to keep up with changes in video speed (e.g., through a time warp or motion smoothing). The results in audio that is unnatural, distorted, or out of synchronization with the visual content, thereby compromising the overall experience of the user.
The mismatch between the rendered VR image and the viewport of the user at the time of display leads to motion sickness. Techniques like motion smoothing or time warp, which adjust the video frame rate to reduce motion sickness and jitter, often degrade audio quality.
The eyes receive information from the virtual environment almost instantly, forming a perception of orientation, movement, and space. However, if the visual content is out of synchronization with the body’s movements, this leads to motion sickness. The conventional systems fail to adjust audio with the same low latency and precision as video, causing a breakdown in the coordination between the image and body movement.
In the conventional virtual reality system, where distortion and unnatural sounds occur frequently when the user moves the head, in accordance with existing arts. VR controllers may be used by the user to interact with the VR environment. The problem arises due to inconsistencies between the rendered VR image at the user’s viewport and the audio rendering. As a result, the user experiences audio that is unnatural, distorted, or unsynchronized, which may compromise the immersive experience. When the user moves their head, a VR engine updates a visual scene to match the new viewpoint. However, an audio engine may not update with the same speed and accuracy, leading to inconsistencies. Further, due to the lag or mismatch in audio rendering, the sounds may become distorted or fail to match the visual cues accurately. This creates a jarring and unnatural experience for the user. If the audio engine does not keep up with rapid changes as detected by the VR engine, the result i.e., VR content may include unsynchronized audio that does not match the visual movements and actions within the VR environment. The environment highlights the challenges faced in the conventional VR systems where the synchronization between the visual and auditory elements is not adequately managed. These inconsistencies result in an overall degraded VR experience, with the users experiencing unnatural, distorted, or unsynchronized audio, especially during head movements.
Based on the current scene, the position and direction of both audio and video are determined. Motion sickness often occurs when there is unsynchronized audio and video in the VR content. In conventional video rendering of VR systems, only the video speed is adjusted, while the audio speed remains unchanged throughout the entire scene. This mismatch leads to a less immersive environment in a conventional VR effect, as the audio does not adapt to the changes in visual cues, failing to provide a cohesive and immersive experience. Inconsistency of a rendered VR image and the user’s viewport at the time of scan-out causes Motion Sickness. Distortion and unnatural sounds are often associated when the user moves the head.
Therefore, there is a need to provide techniques to overcome the above-mentioned challenges in the VR environment.
This summary is provided to introduce methods, in a simplified format, that are further described in the detailed description of the inventive concepts. This summary is neither intended to identify key or essential inventive concepts nor is it intended for determining the scope of the inventive concepts. Embodiments provide techniques that overcome the above-discussed challenges related to enhanced audio rendering for an immersive virtual reality experience.
In an embodiment, an electronic apparatus comprising: a memory, at least one processor comprising a processing circuit, wherein the at least one processor configured to obtain a content associated with a virtual reality environment, wherein the content includes audio data and image data, obtain video frame information associated with the image data, wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data, obtain audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics.
The audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
The at least one processor may obtain an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtain the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
The at least one processor may output the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
The at least one processor may identify a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtain the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
The sound spatial position may indicate a location of the sound source in three-dimensional space relative to a user's position.
The at least one processor may update the sound spatial position in real-time based on a change of the video frame speed.
The at least one processor may obtain the audio spatial characteristics based on one or more spatial audio cues in the audio data, wherein the one or more spatial audio cues may include at least one of increasing or decreasing speed associated with a sound source.
The at least one processor may obtain the audio spatial characteristics based on at least one visual cues in the image data, wherein the at least one visual cues may include information related with an object corresponding to a sound source.
The at least one processor may identify motion sickness degree of a user based on the video frame information associated with the virtual reality environment, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the motion sickness degree.
In an embodiment, an method for controlling an electronic apparatus, the method comprising: obtaining a content associated with a virtual reality environment, wherein the content may include audio data and image data, obtaining video frame information associated with the image data, wherein the video frame information may include a video frame speed for displaying at least one video frame included in the image data, obtaining audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtaining enhanced audio data by rendering the audio data based on the audio spatial characteristics.
The audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
The obtaining the audio spatial characteristics may include obtaining an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtaining the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
The method may include outputting the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
The obtaining enhanced audio data may include identifying a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtaining the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
In embodiments, a method of enhanced audio rendering for an immersive virtual reality experience is disclosed. The method includes receiving a variable video frame information associated with a virtual reality environment, the variable video frame information comprising a speed of one or more variable image frames associated with the virtual reality environment. Further, the method includes generating an audio processing module based on the speed of the one or more variable image frames for preserving one or more spatial audio cues and ensuring that the audio processing module maintains a sense of immersion for a user during time-manipulated variable video frame information being experienced within the virtual reality environment. Furthermore, the method includes determining at least one of localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module, in addition to, rendering the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
In embodiments, a system of enhanced audio rendering for an immersive virtual reality experience is disclosed. The system includes a memory, and at least one processor coupled to the memory. The at least one processor is configured to receive variable video frame information associated with a virtual reality environment, the variable video frame information comprising a speed of one or more variable image frames associated with the virtual reality environment. Further, at least one processor is configured to generate an audio processing module based on the speed of the one or more variable image frames for preserving one or more spatial audio cues and ensuring that the audio processing module maintains a sense of immersion for a user during time-manipulated variable video frame information being experienced within the virtual reality environment. Furthermore, at least one processor is configured to determine at least one of localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module. In addition, at least one processor is configured to render the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
To further clarify the advantages and features of the inventive concepts, a more particular description of the inventive concepts will be rendered by reference to specific examples thereof, which are illustrated in the appended drawings. It is appreciated that these drawings depict only typical examples of the inventive concepts and are therefore not to be considered limiting in its scope. The inventive concepts will be described and explained with additional specificity and detail in the accompanying drawings.
These and other features, aspects, and advantages of the inventive concepts will become better understood when the following detailed description is read with reference to the accompanying drawings in which like characters represent like parts throughout the drawings, wherein:
Figure 1 illustrates a block diagram depicting an embodiment of an audio engine, in accordance with an embodiment of the present disclosure;
Figure 2 illustrates a block diagram of an environment of enhanced audio rendering for an immersive virtual reality experience for a user within a virtual environment, in accordance with an embodiment of the present disclosure;
Figure 3 illustrates a block diagram of a system, in accordance with an embodiment of the present disclosure;
Figure 4 illustrates a flow chart depicting a method of enhanced audio rendering for an immersive virtual reality experience for a user within a virtual environment, in accordance with an embodiment of the present disclosure;
Figure 5 illustrates a flow chart depicting a method for improving voice quality using a vocoder of the system, according to embodiments disclosed herein;
Figure 6A illustrates an example block diagram depicting a noise identification model of a classification module, according to the embodiments disclosed herein;
Figure 6B illustrates an example diagram depicting an embodiment of the classification module, according to embodiments disclosed herein;
Figure 7A illustrates a block diagram depicting an object-based identification model, according to embodiments disclosed herein;
Figure 7B illustrates a block diagram depicting a multimodal analysis of videos of the object-based identification model, according to embodiments disclosed herein;
Figure 7C illustrates another block diagram depicting an object-based identification model, according to embodiments disclosed herein;
Figure 8 illustrates an example diagram depicting the localization and movement module, according to embodiments disclosed herein;
Figure 9 illustrates another example diagram depicting a time-scale modification pipeline of a spatial effect module, according to embodiments disclosed herein;
Figure 10A illustrates an example block diagram depicting a resampling module, according to the embodiments disclosed herein; and
Figure 10B illustrates another example block diagram depicting the resampling module, according to the embodiments disclosed herein.
Figure 11 illustrates an example block diagram control method for the electronic apparatus.
Further, skilled artisans will appreciate that those elements in the drawings are illustrated for simplicity and may not have necessarily been drawn to scale. For example, the flow charts illustrate the method in terms of the most prominent operations involved to help to improve understanding of aspects of the inventive concepts. Furthermore, in terms of the construction of the device, one or more components of the device may have been represented in the drawings by conventional symbols, and the drawings may show only those specific details that are pertinent to understanding embodiments of the inventive concepts so as not to obscure the drawings with details that will be readily apparent to those of ordinary skill in the art having the benefit of the description herein.
For the purpose of promoting an understanding of the principles of the inventive concepts, reference will now be made to the examples illustrated in the drawings and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the inventive concepts is thereby intended, such alterations and further modifications in the illustrated system, and such further applications of the principles of the inventive concepts as illustrated therein being contemplated as would normally occur to one skilled in the art to which the inventive concepts relate.
It will be understood by those skilled in the art that the foregoing general description and the following detailed description are explanatory of the inventive concepts and are not intended to be restrictive thereof.
Reference throughout this specification to “an aspect”, “another aspect” or similar language means that a particular feature, structure, or characteristic described in connection with embodiments is included in at least one embodiment of the inventive concepts. Thus, appearances of the phrase “in an embodiment”, “in one or more embodiments”, “in another embodiment”, and similar language throughout this specification may, but do not necessarily, all refer to the same embodiment.
The terms “comprise”, “comprising”, or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process or method that comprises a list of operations does not include only those operations but may include other operations not expressly listed or inherent to such process or method. Similarly, one or more devices or sub-systems or elements or structures or components preceded by “comprises... a” does not, without more constraints, preclude the existence of other devices or other sub-systems or other elements or other structures or other components or additional devices or additional sub-systems or additional elements or additional structures or additional components.
Embodiments herein and the various features and advantageous details thereof are explained more fully with reference to the non-limiting examples that are illustrated in the accompanying drawings and detailed in the following description. Descriptions of well-known components and processing techniques are omitted so as to not unnecessarily obscure embodiments herein. Also, the various examples described herein are not necessarily mutually exclusive, as some examples may be combined with one or more other examples to form new examples. The term “or” as used herein, refers to a non-exclusive or unless otherwise indicated. The examples used herein are intended merely to facilitate an understanding of ways in which embodiments herein may be practiced and to further enable those skilled in the art to practice embodiments herein. Accordingly, the examples should not be construed as limiting the scope of embodiments herein.
As is traditional in the field, embodiments may be described and illustrated in terms of modules or engines that carry out a described function or functions. These modules or engines, which may be referred to herein as units or blocks or the like, or may include blocks or units, may be physically implemented by analog or digital circuits such as logic gates, integrated circuits, microprocessors, microcontrollers, memory circuits, passive electronic components, active electronic components, optical components, hardwired circuits, or the like, and may optionally be driven by firmware and software. The circuits may, for example, be embodied in one or more semiconductor chips, or on substrate supports such as printed circuit boards and the like. The circuits constituting a block may be implemented by dedicated hardware, by a processor (e.g., one or more programmed microprocessors and associated circuitry), or by a combination of dedicated hardware to perform some functions of the block and a processor to perform other functions of the block. Each block may be physically separated into two or more interacting and discrete blocks without departing from the scope of the inventive concepts. Likewise, the blocks may be physically combined into more complex blocks without departing from the scope of the inventive concepts.
The accompanying drawings are used to help easily understand various technical features and it should be understood that embodiments presented herein are not limited by the accompanying drawings. As such, the present disclosure should be construed to extend to any alterations, equivalents, and substitutes in addition to those which are particularly set out in the accompanying drawings. Although the terms first, second, etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are generally only used to distinguish one element from another.
Embodiments will be described below in detail with reference to the accompanying drawings.
Figure 1 illustrates a block diagram 100 depicting an embodiment of an audio engine, in accordance with an embodiment of the present disclosure. As shown in Figure 1, an audio engine 101 may be communicated with a Virtual Reality (VR) engine 106. The audio engine 101 may include an audio rendering engine 114, and an audio processing module 116. In an embodiment, when a user moves their head, the VR engine 106 updates a visual scene to match the new viewpoint. The audio engine 101 may be adapted to update with the same speed and accuracy. The audio engine 101 dynamically may be adapted to adjust audio playback in real-time to match the visual changes detected by the VR engine 106. In some embodiments, the audio engine 101 may include advanced spatial audio rendering techniques to accurately simulate the direction and distance of sound sources relative to the user’s position and movements. Further, the audio engine 101 may be adapted to maintain constant communication with the VR engine 106, receiving real-time updates about the visual changes and adjusting the audio output accordingly.
In an embodiment, VR controllers 102 may track hand movements and provide input to an operating system 104. The operating system 104 may include a device driver(s) 110 and an original original Software Development Kit (SDK)) 112 adapted to manage hardware resources and provide a platform for VR applications. The device drivers 110 and the original SDK 112 may be configured to facilitate communication between the hardware and software components of the VR engine 106. A VR content 108 may represent digital content and scenes that the user experiences within the virtual environment.
Figure 2 illustrates a block diagram 200 of an environment of enhanced audio rendering for an immersive virtual reality experience for a user 203 within a virtual environment, in accordance with an embodiment of the present disclosure. As shown in Figure 2, the user 203 is using a wearable device 202. For instance, the wearable device 202 is a head-mounted display device configured to display a VR environment. The wearable device 202 may be configured to generate audio and video data to enable the user 203 to experience the VR environment. The wearable device 202 may be connected to a system 204 configured to provide the enhanced audio rendering for the wearable device 202.
The wearable device 202 may be connected to a system 204 configured to provide the enhanced audio rendering for the wearable device. Then explain the system 204 that it may be located remotely or within the wearable device. Further, it should be noted that although Figure 2 depicts the system 204 as separate from the wearable device 202. In one embodiment, the system 204 may be integrated into the wearable device 202. The system 204 may include the VR engine 106, the audio rendering engine 114, the audio processing module 116, a scene change detector 118, a viewport change detector 206, and a point of view prediction 208. The VR engine 106 may be adapted to create and manage the virtual environment in which the user interacts. The audio rendering engine 114 may be adapted to process and generate audio in the VR environment. The audio rendering engine 114 plays a critical role in maintaining audio-visual synchronization and providing an immersive auditory experience. The audio processing module 116 may handle detailed processing and manipulation of audio signals.
In an embodiment, the scene change detector 118 may be adapted to continuously monitor the VR environment for significant changes in the visual scene. The significant changes may include transitions between different scenes, major shifts in visual content, or changes in the VR environment that require corresponding adjustments in audio rendering. For example, the scene change detector 118 identifies when the user 203 moves from one scene to another or when a significant event occurs within the scene, such as entering a new room or experiencing an explosion within the VR environment.
In an embodiment, the viewport change detector 206 may be configured to track changes in the user’s viewport within the VR environment. The track changes may include monitoring the user’s head movements and the corresponding changes in the visible portion of the VR scene. Further, the viewport change detector 206 may be configured to inform the audio engine 101 to modify spatial audio rendering based on the viewpoint, ensuring that sounds are accurately localized and move naturally with the user’s head movements.
In an embodiment, the point of view prediction 208 may be adapted to predict the future head movements of the user 203 and changes in viewpoint. The point of view prediction 208 may help preemptively adjust the audio rendering to maintain synchronization with the visual content.
In an embodiment, the wearable device 202 may be configured to receive variable video frame information associated with the virtual reality environment. The variable video frame information may include information such as, but not limited to, a speed of variable image frames associated with the virtual reality environment. The system 204 may be configured to may be configured to generate the audio processing module 116 based on the speed of the variable image frames for preserving spatial audio cues and ensuring that the audio processing module 116 maintains the sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment. The audio processing module 116 may include the audio features based on the received variable video frame information. The spatial audio cues provide information about the location and the movement of sound sources in the virtual environment, the spatial audio cues include speed up or speed down at which sound is coming.
In an embodiment, the system 204 may be configured to determine localization and movement of sounds in the virtual environment based on the speed of the variable image frames, and the generated audio processing module 116. In an embodiment, the movement of sounds may be simulated based on the speed of the variable image frames associated with the virtual reality environment and the determined localization and movement of sounds. Further, the system 204 may be configured to determine the movement of sounds based on the speed of variable image frames associated with the virtual reality environment and changes in visual cues, the visual cues within a current scene provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds.
In some embodiments, the system 204 may be configured to render the enhanced audio in the VR environment for the immersive virtual reality experience based on the determined localization and movement of sounds. In an embodiment, render the enhanced audio may include rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects. The rendering enhanced audio may provide synchronization of audio with respect to the speed of the variable image frames associated with the virtual reality environment.
Further, the system 204 may be configured to identify spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined localization and movement of sounds. The spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203. Furthermore, the system 204 may be configured to update the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment.
In an embodiment, the system 204 may be configured to predict motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment. Further, the system 204 may be configured to determine the movement of the user 203 based on the variable video frame information associated with the virtual reality environment.
In an embodiment, Further, the system 204 may include extracting features corresponding to the variable video frame information for noise, speech, music, and object tagging. Furthermore, the system 204 may include a plurality of audio features associated with the speed of the variable image frames.
In an embodiment, the VR content 108 may represent digital content and scenes that the user 203 experiences within the virtual environment. The VR content 108 may include both visual and audio elements to create an immersive experience. The scene change detector 118 may be responsible for monitoring the current VR scene and determining any changes or transitions. For example, the scene change detector 118 may identify when the visual content changes significantly, such as moving to a new scene or a significant event occurring within the scene. Further, the scene change detector 118 may predict the user’s future head movements and adjust the audio accordingly to maintain synchronization.
In an embodiment, uses data from the scene change detector 118 to assess the potential for motion sickness and make necessary adjustments to the audio rendering to mitigate it. Further, for video speed change, the system 204 detects any changes in the speed of the video content based on the predicted scene changes. The system 204 may monitor and detect changes in video playback speed that may be required to keep the audio and visual elements in synchronization. Sound process, the system 204 may be adapted to determine localization and movement of sound based on changes in video speed. Audio rendering, the system 204 may provide time-manipulated audio based on user orientation and speed of visual cues. Finally, immersive content, the system 204 may provide enhanced immersive audio to reduce the motion sickness of the user 203 in the current scene.
Figure 3 illustrates a block diagram of system 204, in accordance with an embodiment of the present disclosure. Figure 4 illustrates a flow diagram depicting a method 400 of enhanced audio rendering for an immersive virtual reality experience for the user 203 within a virtual environment, in accordance with an embodiment of the present disclosure. For the sake of brevity, the description of Figures 2, 3, and 4 are explained in conjunction with each other.
Referring to Figure 3, the system 204 may include, but is not limited to, a processor 304, memory 302, an interface 306, and a plurality of modules 308. The memory 302, the interface 306, and the plurality of modules 308 may be coupled to the processor 304.
The processor 304 may be a single processing unit or several units, all of which could include multiple computing units. The processor 304 may be implemented as one or more microprocessors, microcomputers, microcontrollers, digital signal processors, central processing units, state machines, logic circuitries, and/or any device that manipulates signals based on operational instructions. Among other capabilities, the processor 304 is configured to fetch and execute computer-readable instructions and data stored in the memory 302.
The memory 302 may include any non-transitory computer-readable medium known in the art including, for example, volatile memory, such as static random access memory (SRAM) and dynamic random access memory (DRAM), and/or non-volatile memory, such as read-only memory (ROM), erasable programmable ROM, flash memories, hard disks, optical disks, and magnetic tapes. Further, the memory 302 may include an operating system 312 for performing one or more tasks of the system 204, as performed by a generic operating system 312 in the communications domain.
In an embodiment, the processor 304 may be configured to receive variable video frame information associated with a virtual reality environment. The variable video frame information may include the speed of variable image frames associated with the virtual reality environment. The processor 304 may be configured to generate the audio processing module 116 based on the speed of the variable image frames for preserving the spatial audio cues and ensuring that the audio processing module 116 maintains a sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment. The spatial audio cues may provide information about the location and the movement of sound sources in the virtual environment. The spatial audio cues may include the speed up or the speed down at which sound is coming.
Further, the processor 304 may be configured to determine the localization and movement of sounds in the virtual environment based on the speed of the variable image frames, and the generated audio processing module 116. Furthermore, the processor 304 may be configured to render the enhanced audio in the virtual environment for the immersive virtual reality experience based on the determined localization and movement of sounds. The movement of sounds may be simulated based on the speed of the variable image frames associated with the virtual reality environment and the determined localization and movement of sounds. The rendering enhanced audio may include rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects, the rendering enhanced audio may provide the synchronization of audio with respect to the speed of the variable image frames associated with the virtual reality environment.
In some embodiments, the processor 304 may be configured to identify spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined at least one of localization and movement of sounds, wherein the spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203. The processor 304 may be configured to update the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment. The audio processing module 116 may include the plurality of audio features based on the received variable video frame information. The plurality of audio features may be associated with the speed of the variable image frames.
In an embodiment, the processor 304 may be configured to determine the movement of sounds based on the speed of the variable image frames associated with the virtual reality environment and the changes in visual cues. The visual cues within the current scene may provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds. Further, the processor 304 may be configured to predict the motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment. The processor 304 may be configured to determine the movement of the user 203 based on the variable video frame information associated with the virtual reality environment. Further, the processor 304 may be configured to extract features corresponding to the variable video frame information for noise, speech, music, and object tagging.
The plurality of modules 308 amongst other things, include routines, programs, objects, components, data structures, etc., which perform particular tasks or implement data types. The plurality of modules 308 may also be implemented as, signal processor(s), state machine(s), logic circuitries, and/or any other device or component that manipulates signals based on operational instructions. The plurality of modules 308 may be configured to perform the steps of the present disclosure using the data stored in a database 310 for automated parking management, as discussed herein. In one embodiment, the database 310 may be configured to store the information as required by the plurality of modules 308 and the one or more processors 304 for enhanced audio rendering for the immersive virtual reality experience for the user 203 within the virtual environment.
Further, the plurality of modules 308 can be implemented in hardware, instructions executed by a processing unit, or by a combination thereof. The processing unit can comprise a computer, a processor, such as the processor 304, a state machine, a logic array, or any other suitable wearable device capable of processing instructions. The processing unit can be a general-purpose processor which executes instructions to cause the general-purpose processor to perform the required tasks or, the processing unit can be dedicated to performing the required functions. In another embodiment of the present disclosure, the plurality of modules 308 may be machine-readable instructions (software) which, when executed by a processor/processing unit, perform any of the described functionalities.
In some embodiments, the plurality of modules 308 may include a set of instructions that may be executed to cause the system 204 to perform any one or more of the methods disclosed herein. The plurality of modules 308 may be configured to perform the steps of the present disclosure using the data stored in the memory 302 to provide enhanced audio rendering for the immersive virtual reality experience for the user 203 within the virtual environment, as discussed throughout this disclosure. In an embodiment, each of the modules 308 may be hardware units that may be outside the memory 302.
In an embodiment, the plurality of modules 308 may include a detection module 310, a prediction module 313, the audio rendering engine 114, and the audio processing module 116. The audio rendering engine 114 may include a spatial effect module 314, a resampling module 316, and a vocoder 318. Further, the audio processing module 116 may include sub-modules such as an audio feature extraction module 320, a classification module 322, and a localization and movement module 324. The plurality of modules 308 and its working is further explained in detail with reference to Figure 3B.
In an embodiment, the detection module 310 may be configured to identify changes and events within the VR environment that require adjustments in audio and visual rendering. The detection module 310 may be configured to continuously monitor the VR scene and detect user interactions, head movements, and scene transitions. For example, the detection module 310 monitors and identifies changes in the user’s viewpoint within the VR environment. The detection module 310 continuously tracks the head movements of the user 203 to determine changes in the viewing angle and direction. The detection module 310 may detect when the user’s viewport (for example, the visible area of the virtual environment) shifts due to head movements or other inputs.
Further, the detection module 310 may be configured to identify significant changes within the VR environment. The detection module 310 may be configured to detect specific events or changes that indicate a scene transition, such as entering a new room, an explosion, or the appearance of new objects.
The prediction module 313 may be configured to anticipate future movements and actions of the user 203 based on current and past data. By analyzing patterns in the user's behavior and movements, the prediction module 313 may be configured to predict the user’s next position or action. The prediction module 313 may be adapted to pre-render images and pre-process audio, reducing latency and ensuring smoother transitions and more accurate synchronization between the audio and visual elements.
The spatial effect module 314 may be configured to apply spatial effects to the audio signals, creating a 3D soundscape that matches the virtual environment. The spatial effect module 314 may adjust the direction, distance, and movement of sounds to align with the user’s perspective and the visual scene. The resampling module 316 may be configured to adjust the sampling rate of audio signals to ensure they match the current playback speed and synchronization requirements.
The vocoder 318 may process and manipulate the audio signals to alter pitch and timing to improve the quality. The vocoder 318 ensures that voice and other audio effects remain natural and intelligible, even when there are changes in playback speed. The Vocoder 318 may be a type of vocoder-purposed algorithm which is used to interpolate information present in the frequency and time domains of audio signals by using phase information extracted from a frequency transform.
The audio feature extraction module 320 may be configured to extract acoustic features from audio signals (for example, a speech signal), such as pitch, tone, and amplitude. The features may be used for further processing and analysis thereby ensuring that the audio remains high-quality and immersive. The acoustic features may include Linear Predictive Cepstral Coefficients (LPCC) features and Mel-Frequency Cepstral Coefficients (MFCC) features. The LPCC features may include, but are not limited to, 13 delta LPCC features, 13 delta-delta LPCC features, 13 LPCC features, and the like. The MFCC features may include 12 MFCC Cepstral Coefficients, 12 Delta MFCC features, 12 Double Delta MFCC features, 1 energy coefficient, 1 delta energy coefficient, 1 double Delta energy coefficient, and the like. In an embodiment, Combining LPCC and MFCC features may provide a more comprehensive representation of the speech signal. By applying Principal Component Analysis (PCA) to the combined features, dimensionality may be reduced while preserving the important information. The combined use of LPCC and MFCC features, along with the PCA, may enhance the capability of the system 204 to analyze and process speech signals within the VR environment.
The classification module 322 may be configured to categorize the audio signals based on features and context within the VR environment. The classification module 322 may be configured to identify different types of sounds such as dialogue, ambient noise, and sound effects to apply processing techniques.
The localization and movement module 324 may be configured to determine spatial positions and movements of sound sources within the virtual environment. The localization and movement module 324 may ensure that audio cues are accurately positioned and move in harmony with the visual scene, enhancing the sense of presence and immersion for the user 203.
The modules 310, 312, 114, and 116 may be in communication with each other. In an embodiment, the modules 310, 312, 114, and 116 may be a part of the processor 304. In another embodiment, the processor 304 may be configured to perform the functions of the modules 310, 312, 114, 116.
At least one of the modules 310, 312, 114, and 116 may be implemented through an artificial intelligence (AI) model. A function associated with AI model may be performed through the non-volatile memory, the volatile memory, and the processor 304. Accordingly, the processor 304 may include one or a plurality of processors. At this time, one or a plurality of processors may be a general purpose processor, such as a central processing unit (CPU), an application processor (AP), or the like, a graphics-only processing unit such as a graphics processing unit (GPU), a visual processing unit (VPU), and/or an AI-dedicated processor such as a neural processing unit (NPU). The one or a plurality of processors control the processing of the input data in accordance with a predefined operating rule or artificial intelligence (AI) model stored in the non-volatile memory and the volatile memory. The predefined operating rule or artificial intelligence model is provided through training or learning. The learning may be performed in a device in which AI model according to an embodiment is performed, and/or may be implemented through a separate server/system.
The AI model may include a plurality of neural network layers. Each layer has a plurality of weight values, and performs a layer operation through calculation of a previous layer and an operation of a plurality of weights. Examples of neural networks include, but are not limited to, convolutional neural network (CNN), deep neural network (DNN), recurrent neural network (RNN), restricted Boltzmann Machine (RBM), deep belief network (DBN), bidirectional recurrent deep neural network (BRDNN), generative adversarial networks (GAN), and deep Q-networks.
The learning technique is a method for training a predetermined target device (for example, a robot) using a plurality of learning data to cause, allow, or control the target device to make a determination or prediction. Examples of learning techniques include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
It should be noted that the system 204 may be a part of the wearable device 202. In another embodiment, the system 204 may be connected to the wearable device 202. In such embodiment, the wearable device 202 may be a virtual reality (VR) headset designed to provide the immersive VR experience. The wearable device 202 may include integrated audio and visual systems that allow the user 203 to interact with and perceive the virtual environment.
Figure 4 illustrates the flow chart depicting a method 400 of the enhanced audio rendering for an immersive virtual reality experience for the user within a virtual environment, according to embodiments disclosed herein. The method 400 may be performed at the wearable device 202. The method 400 may be performed by the system 204 comprising the processor 304 and the memory 302.
At block 402, the method 400 includes receiving, at the wearable device 202, variable video frame information associated with the virtual reality environment, the variable video frame information may include information such as, but not limited to, a speed of variable image frames associated with the virtual reality environment.
At block 404, the method 400 includes generating the audio processing module 116 based on the speed of the variable image frames for preserving spatial audio cues and ensuring that the audio processing module 116 maintains the sense of immersion for the user 203 during time-manipulated variable video frame information being experienced within the virtual reality environment. The audio processing module 116 includes the audio features based on the received variable video frame information. The spatial audio cues provide information about the location and the movement of sound sources in the virtual environment, the spatial audio cues include speed up or speed down at which sound is coming.
At block 406, the method 400 includes determining the localization and movement of sounds in the virtual environment based on the speed of the one or more variable image frames, and the generated audio processing module 116. In an embodiment, the movement of sounds is simulated based on the speed of the one or more variable image frames associated with the virtual reality environment and the determined localization and movement of sounds. Further, determining the movement of sounds based on the speed of variable image frames associated with the virtual reality environment and changes in visual cues, the visual cues within a current scene provide information about the virtual environment and help the user 203 make sense of spatial relationships, motion, and positioning of objects and sounds.
At block 408, the method 400 includes rendering the enhanced audio in the VR environment for the immersive virtual reality experience based on the determined localization and movement of sounds.
In an embodiment of present disclosure, the method 400 includes identifying spatial positions of sounds within the virtual environment based on the speed of the variable image frames and the determined localization and movement of sounds. The spatial positions of sounds indicate three-dimensional coordinates that define the location of sound sources relative to the position of the user 203. The method includes updating the spatial positions in real-time as speed changes in the variable image frames associated with the virtual reality environment.
In an embodiment, the method 400 includes predicting motion sickness of the user 203 based on the variable video frame information associated with the virtual reality environment. Further, the method 400 includes determining the movement of the user 203 based on the variable video frame information associated with the virtual reality environment.
In an embodiment, the method 400 includes rendering the enhanced audio includes rendering to synchronize with the visual cues, resampling audio signals, and applying spatial effects. The rendering enhanced audio provides synchronization of audio with respect to the speed of the one or more variable image frames associated with the virtual reality environment. Further, the method 400 includes extracting features corresponding to the variable video frame information for noise, speech, music, and object tagging. Furthermore, the method 400 includes the plurality of audio features associated with the speed of the variable image frames.
In an embodiment of the present disclosure, the audio feature extraction module 320 may be configured to process the audio signal to extract the features for further processing. The steps involve several transformations and computations, which are detailed below.
Figure 5 illustrates the flow chart depicting a method 500 for improving voice quality using the vocoder of the system 204, according to embodiments disclosed herein. The method 500 may be performed by the system 204 comprising the processor 304 and the memory 302 of the system 204.
At block 502, the method 500 includes accepting the raw audio waveform (speech) as the initial input.
At block 504, the method 500 includes transforming the audio waveform into a mel-spectrogram representation to capture both frequency and time information.
At block 506, the method 500 includes normalizing the Mel-spectrogram to ensure consistent data for further processing.
At block 508, the method 500 includes converting the normalized Mel-spectrogram into the format required by the WaveNet decoder.
At block 510, the method 500 includes generating audio samples one by one in a casual autoregressive manner, using previously generated samples and mel-spectrogram information to predict the next sample.
At block 512, the method 500 includes applying an inverse filter to the generated audio to remove any pre-emphasis applied during preprocessing, ensuring a more natural sound.
At block 514, the method 500 includes adjusting the audio signal to an appropriate volume level to ensure consistency and enhance the listening experience.
At block 516, the method 500 includes delivering the improved speech audio waveform with enhanced quality.
The various operations of methods described above may be performed by any suitable device capable of performing the operations, such as the processing circuitry discussed above. For example, as discussed above, the operations of methods described above may be performed by various hardware and/or software implemented in some form of hardware (e.g., processor, ASIC, etc.).
Figure 6A illustrates an example block diagram 600 depicting a noise identification model of the classification module 322, according to the embodiments disclosed herein. The block diagram 600 includes a Residual Neural Network (ResNet) module 602, a Visual Geometry Group VGGish (for example, CNN Algorithm) module 604, a Neural Network (NN) classification module 606, and noise segments module 608. In an embodiment, the input 610a may be a single video, and the output may be the noise segments 608. The noise segments 608 may include an identification of noise or not or corresponding sound.
In an embodiment, the ResNet module 602 may be adapted to process video frames to extract deep image features. The deep image features may capture the spatial and temporal context of the video. The VGGish module 604 may be required for audio classification tasks. The VGGish module 1004 may be adapted to convert the audio features into a format suitable for neural network processing. Further, the NN classification module 606 may be configured to take combined features from the ResNet module 602 and the VGGish module 604. The NN classification module 606 may be configured to perform the final classification to determine if the segment is “noise” or “not noise”. In an embodiment, the noise segments 608 may be the final output consists of the classification results for each segment. Each segment may include noise and not noise. The noise may indicate the presence of noise and not noise may indicate an absence of noise corresponding to sound.
Figure 6B illustrates an example diagram depicting an embodiment of the classification module 322, according to embodiments disclosed herein. The classification module 322 may include the audio signal 610b, a music classifier 612, a song extractor 614, and speech, and music sequences 616. The audio signal 610b, the music classifier 612, the song extractor 614, and the speech, and music sequences 616 are also part of the modules 208 of the system 204. The audio signal 610b may be extracted from the video 610a, which serves as the raw data for the classification. The music classifier 612 may include a classification model, specifically a Support Vector Machine (SVM) classifier, which processes the audio signal 610b to distinguish between music and non-music segments. The music and the non-music segments of video m may include segmenting video m into frames with length l, calculating audio features for each frame, classifying video frames in music and non-music classes using a Support Vector Machine, and returning music and non-music segments.
The song extractor 614 may be a component that processes segments identified as music to extract song sequences, leveraging the fact that songs are typically accompanied by music. The song extractor 614 may be configured to take music component provided by the music classifier 612 as an input and searches for song sequences in it. The music part certainly includes the songs, but there may also be dialogue and action scenes with background music in the music component. To filter out dialogue and action scenes, the song extractor 614 may be divided into three sub-components such as a potential song sequence generator 614a, a vocal detector 614b, and a vocal pattern checker 614c. The potential song sequence generator 614a may be configured to accept the music part of a video and generate a number of potential song sequences (PSS), which are highly likely to be songs. The potential song sequence generator 614a may break the music part of the video 610a into multiple sequences where a sequence refers to consecutive musical seconds without break. The potential song sequence generator 614a then applies rules of sequences in order to extract a potential song sequence. The vocal detector 614b may be configured to detect vocal and non-vocal segments of the possible song sequence. The vocal detector 614b may be configured to use the features for music or non-music classification spectral centroid and mel-frequency cepstral coefficients features. Similar to that of the music classifier 612 along with some additional audio features, spectral centroid and Mel-frequency cepstral coefficients may be used. The spectral centroid may be defined as a weighted mean of signal frequencies. The spectral centroid may be determined by a Fourier transform using frequency magnitudes as weights.
The vocal pattern checker 614c may be configured to check the pattern of vocal and non-vocal segments of possible song sequences detected by the vocal detector 614b. The vocal pattern checker 614c may use the pattern of vocal and non-vocal segments to differentiate between Potential Song Sequences (PSS) that are songs and the PSS that are action or dialogues. For example, if a vocal segment appears in the song, the vocal segment may be either intro, chorus, verse, or outro. The duration for vocal segments may fall in a range between 20 to 100 seconds. On the other hand, if a non-vocal segment appears in the song, the non-vocal segment may be a bridge. Depending on the place of its appearance, the duration of the bridge in non-vocal segments ranges from 20 to 40 seconds. The pattern of vocal and non-vocal segments may be used to differentiate between potential song sequences which are songs to the PSS which are action or dialogues.
The performance of the classification of music from audio depends mainly on the selection of appropriate audio features and classifiers. The audio features may include a zero-crossing rate, an audio spectrum flux, a short-time energy or sub-band energy distribution, and an audio intensity. Zero Crossing Rate may be defined as measures the rate at which the audio signal changes sign, indicating the signal’s noisiness or the presence of high frequencies. The audio spectrum flux, may be defined ascaptures the rate of change in the power spectrum of the audio signal, indicating the presence of dynamic elements such as music or speech. The short-time Energy or sub-band energy Distribution may be defined as analyzes the energy content within short windows of the audio signal, identifying variations characteristic of music or speech. The audio intensity, measures the perceived loudness of the audio signal, helping to distinguish between music and quieter non-music segments.
Further, the output of the song extractor 614 may categorize the input audio signal into distinct speech and music sequences 616 for further analysis or processing. The speech and music sequences 616 may provide a clear categorization of the audio signal into speech and music sequences. The speech and music sequences 616 may be used for various applications such as content indexing, music retrieval, and audio analysis.
Figure 7A illustrates a block diagram 700a depicting an object-based identification model, according to embodiments disclosed herein. The object-based identification model 702 may be configured to identify structural and semantic relationships between a non-subject or subject audio and the video frames by object identification segments (for example, frame objects tagger) 704. The object-based identification model 702 may include a segmentation module 702a adapted to extract the objects for tagging. Further, the object-based identification model 702 may include a Motion Scale Invariant feature Transform (Mo SIFT) 702b, a Mel-Frequency Cepstral co-efficient 702c, a Bag of Words (BoW) 702d, a SVM classification (for example, Dynamic Time Wrapping (DTW)) 702e, and another SVM classification 702f.
The object-based identification model 702 may train three two-class SVMs in order to learn tagging models. The first SVM model 702e may be constructed using low-level audio features, and the SVM classification 702f using low-level visual features. Fusing the predictions of the two SVM models i.e., a visual classifier and an audio classifier using another two-class SVM for the final prediction of object identification segments for Frame objects Tagger 704.
The Mo SIFT 702b may be a standard SIFT algorithm applied to find visually distinctive interest points in the spatial domain. The Mo SIFT 702b (for example, 256 dimensions) may be designed to represent the feature point in two parts: the first 128 dimensions are standard SIFT features and the remaining 128 dimensions are the aggregated histogram of optical flow.
In a learning phase, the vocabulary of visual words may be typically defined as the cluster centers obtained from k-means clustering over a large collection of sample low-level descriptors (Mo SIFT descriptors). The BoW 702d represents each video sequence as a histogram over a set of visual words (motion and color key points) to generate a fixed-dimensional encoding that may be processed using the standard classifier.
Figure 7B illustrates a block diagram 700b depicting a multimodal analysis of videos of the object-based identification model, according to embodiments disclosed herein. The object-based identification model 700b includes a visual feature extraction 706a, an audio feature extraction 706b, the visual classifier 706c, the audio classifier 706d, and a multimodal classifier 706e. The visual feature extraction 706a may be the process of analyzing visual data from the video to identify key characteristics and patterns. The characteristics and patterns may include processing frames to extract information such as color, texture, shape, and motion.
The audio feature extraction 706b may include analyzing the audio signal to identify important characteristics such as pitch, rhythm, timbre, and zero crossing rate. The audio feature extraction 706b may convert raw audio data into a set of meaningful features that may be used for further analysis and classification. In an embodiment, the audio feature extraction 706b may include Mel Frequency Cepstral Coeficients (MFCC) that divide each utterance into overlapping frames, each frame is 25ms long with a step of 10ms.
The visual classifier 706c may be an algorithm or model that uses extracted visual features to categorize video content into different classes. The visual classifier 706c may use machine learning techniques like convolutional neural networks (CNNs) to recognize and classify objects, scenes, or actions within the video. In an embodiment, histograms obtained from BoW are high-dimensional vectors that may be classified based on different scenes using a standard classifier, typically the SVM model.
The audio classifier 706d may use extracted audio features to categorize the audio content into different classes. The audio classifier 706d may involve identifying whether a segment contains music, speech, or other types of sounds. In an embodiment, the audio classifier 706d may include Dynamic Time Warping (DTW) that matches two signals by calculating the best path that is the best alignment between the two signals resulting in the level of similarity between them. The SVM may be used to classify the sound using DTW as its distance metrics.
The multimodal classifier 706e may be adapted to combine both visual and audio features to make more accurate and robust classifications. By integrating data from both modalities, the multimodal classifier 706e may understand the context and content of the video, leading to improved performance in tasks such as identifying scenes, actions, or specific events. In an embodiment, both the outputs of visual classifier 706c and the audio classifier 706d may be fed to the multimodal classifier 706e. The final output from the multimodal classifier 706e may be whether the video clip has classification and objects tagging.
Figure 7C illustrates another block diagram 700c depicting an object-based identification model, according to embodiments disclosed herein. The object-based identification model 700c may use existing networks as building blocks which act as a strong core of model and developing a panoptic head that combines outputs of semantic and instance segmentations. The object-based identification model 700c may include input frame 708a, backbone network 708b, semantic head 708c, Instance Head 708d, a Panoptic head 708e, panoptic logits 708f, object tagging 708g, semantic logits 710a, stuff logits 710b, the panoptic logits 708f. Further, Instance logits 710c, occlusion head 710d, an occlusion matrix 710e, fusion 710f, a ground truth mask 710g, and an edge map head 710h. The input frame 708a may be a single image or frame from the video that is fed into the object-based identification model for analysis. The backbone network 708b may be a pre-trained neural network that acts as the core feature extractor. The backbone network 708b may process the input frame 708a to produce high-level feature maps. The backbone networks 708b may include ResNet, VGG, and EfficientNet. Further, the semantic Head 708c may process the feature maps from the backbone network 708b to perform semantic segmentation, which classifies each pixel in the frame into different classes (e.g., road, sky, building). The instance head 708d may process the feature maps to perform instance segmentation, which identifies and delineates individual objects within the frame (e.g., a specific car, a particular person). The panoptic head 708e may combine the outputs from the semantic head 708c and the instance head 708d to create a unified segmentation that includes both “stuff” (background regions like sky, road) and “things” (discrete objects like cars, people). The results in panoptic logits 708f. The panoptic logits 708f may be the output probabilities from the panoptic head 708e, indicating the likelihood of each pixel belonging to a specific class or instance. The object tagging 708f may involve assigning labels or tags to the identified objects based on the panoptic logits. The object tagging 708f may annotate the objects with respective categories (e.g., car, person).
In an embodiment, the semantic logits 710a may be the output probabilities from the semantic head 708c, indicating the likelihood of each pixel belonging to a specific semantic class. The Stuff logits 710b may be the specific probabilities from the semantic logits that correspond to background regions or “stuff” classes (e.g., sky, grass). Further, instance logits 710c may be the output probabilities from the instance head 708d, indicating the likelihood of each pixel belonging to a specific instance of an object. The output probabilities may include instance i, and instance j. The occlusion head 710d may identify and process occlusions within the frame, where one object may partially or fully block another. This occlusion head 710d may help in accurately segmenting overlapping objects. Furthermore, the occlusion matrix 710f may be a representation of occlusions detected by the occlusion head 710d, showing which parts of objects are occluded and how they overlap with each other. The fusion 710f may refer to the process of combining information from various logits (for example, semantic, instance, occlusion) to produce a coherent and accurate final output. The ground truth mask 710g may be the actual annotated data used for training and validating the model. Finally, the edge map head 710h may detect and highlight the edges of objects within the frame, helping to delineate boundaries more precisely.
Figure 8 illustrates an example diagram 800 depicting the localization and movement module, according to embodiments disclosed herein. The localization and movement module 324 may include building audio datasets 802, building interaural phase difference (IPD) datasets 804, and a training model 806. The building audio datasets 802 may include a multichannel impulse response database 802a, a real Room Impulse Response (RIR) 802b, and a simulated Room Impulse Response (RIR) 802c. The building (IPD) interaural phase difference datasets 804 may include short-time Fourier transform (STFT) 804a, IPD feature extraction 804b, real IPD 804c, and a simulated IPD 804d. Further, the training model 806 includes a convolutional neural network plus regression model 806a. Pyroomacoustics platform and the multichannel impulse response database (MIRD) may be used to generate both simulated and real room impulse response (RIR) datasets 804c, 804d.
In an embodiment, the multichannel impulse response database 802a may be a collection of impulse responses recorded from multiple channels, providing a comprehensive set of data for different environments and setups. The real RIR 802b may be recorded in real physical rooms, capturing the acoustics and reflections of actual environments. The simulated RIR 802c may be generated through simulations, allowing for the creation of data that may not be easily captured in real-world settings.
In an embodiment, the STFT 804a may be a mathematical technique applied to audio signals to convert them from the time domain to the frequency domain, facilitating the analysis of frequency components over short time windows. The IPD feature extraction 804b may be the process of extracting interaural phase difference features, which are critical for determining the spatial characteristics of sounds based on the phase differences between signals received at the left and right ears. The real IPD 804c may be derived from real-world recordings, providing authentic spatial audio features. The simulated IPD 804d may be generated through simulations, and used to augment the dataset with controlled variations and scenarios. The IPD features of the sound signal may be firstly extracted from time-frequency domain by the STFT 804a. Then, the IPD features map may be fed to the CNN-R model 806 as an image for sound source localization.
In an embodiment, the machine learning model 806a combines convolutional neural networks (CNNs) with regression techniques. The CNN part may be used for extracting spatial features from the audio data, while the regression part predicts the precise localization and movement parameters of the sound sources. The CNN with a regression model (CNN-R) to estimate the sound source angle and distance based on the acoustic characteristics of the interaural phase difference (IPD).
Figure 9 illustrates another example diagram 900 depicting a time-scale modification pipeline of the spatial effect module 314, according to embodiments disclosed herein. The example diagram 900 may include detailed steps and processes involved in modifying the time scale of an audio signal to maintain its frequency content and ensure high-quality audio output. The example diagram 900 may include the input audio 610b, STN Decomposition 906, a phase vocoder 908, Stretch Sines and Transients, Extract CQT (Constant-Q Transform) Spectrogram, Train WaveNet 912, and envelop detection 910a, 910b, and time stretched output signal 912.
In the STN Decomposition stage 906, the input audio signal may be decomposed as the STN, sines (tonal content), transients (impulsive events), and noise (sound nuances). The sines represent the tonal content of the audio, and transients represent impulsive events or sudden changes in the audio, the noise Represents the background sound nuances. The sines may be processed via a phase vocoder 908 with an identity phase locking. The phase vocoder 908 may be a type of vocoder-purposed algorithm which is used to interpolate information present in the frequency and time domains of audio signals by using phase information extracted from a frequency transform. The vocoder-purposed algorithm may allow frequency-domain modifications to a digital sound file.
The noise component may be neutrally resynthesized at the desired TSM factor. Transients may be preserved and relocated onto the new time axis according to detected peaks, which are isolated due to the nature of the STN decomposition. The envelope detection 910 may be configured to reshape the STN according to the envelope of the original signal to compensate for the pre-echo effect. The neural network (WaveNet) 912 may be trained using a CQT spectrogram and the desired TSM factor. The training may enable the network to generate new noise components that match the modified time scale. Using the trained WaveNet, new noise components may be synthesized to align with the stretched or compressed audio signal, maintaining the original spectral qualities. The processed sines, transients, and the synthesized noise may be combined to form a cohesive audio signal. A final processed audio signal 914 may be delivered, which has been time-stretched or compressed while preserving its frequency content and maintaining high audio quality.
Figure 10A illustrates an example block diagram 1000 depicting a resampling module 316, according to the embodiments disclosed herein. The example block diagram 1000 may include three blocks 1002, 1004, and 1006. At block 1002, the audio model may represent the original audio signal and its properties before any adjustments are made. The audio model may include the raw audio data that needs to be adjusted in response to changes in the video speed. At block 1004, the resampling module 316 may be configured for adaptive sample rate which aims to adjust the audio playback speed to match the modified video speed while preserving pitch and timbre. At block 1006, the resampling module 316 may include the output audio signals that have been resampled to match the new video speed. The output audio signals may be the final product of the resampling module 316, ready to be synchronized with the modified video playback.
Figure 10B illustrates another example block diagram depicting the resampling module, according to the embodiments disclosed herein. In an embodiment, audio sampling is the process of transforming a musical source into a digital file. The more samples you take - known as the ‘sample rate’ - the more closely the final digital file may resemble the original. A higher sample rate tends to deliver a better-quality audio reproduction.
Sample rates may be usually measured per second, using kilohertz (kHz) or cycles per second. The CDs are usually recorded at 44.1kHz - which means that every second, 44,100 samples were taken. Common sample rates are 44.1khz, 48khz, 88.2khz, 96khz and so on.
The resampling module 316 starts by checking a video frame rate 1008. For frame rates 1008 less than 120 frames per second (fps), the audio sample rate is set to 48 kHz 1010. Further, for frame rates higher than 120 fps, pitch changes 1012 may be considered to maintain synchronization. If pitch changes are instantaneous, the resampling module 316 adjusts accordingly. For non-instantaneous changes, the audio latency 1016 may be considered. Medium latency scenarios, for example, speech utterances 1018 use a sample rate of 96 kHz, while high latency scenarios i.e., number of audio sources 1020 use 192 kHz to preserve audio quality.
Background noise levels i.e., ambient environment noise 1014 may be considered to ensure that the chosen sample rate provides clear and accurate audio reproduction. Further, the resampling module 316 adapts the sample rate based on the user’s scenario and the audio artifacts affected. Further, the resampling module 316 may include a decision tree-based approach implemented to uniquely identify and adjust each user where multiple rule-based classifiers are used to learn (dataset). The outputs of the decision tree may be then used to predict the correct output.
In an embodiment, a training phase of the vocoder includes the speech signal, the feature extraction block, the time resolution adjustment block, and a mapping technique block. Further, the testing phase may include the speech signal, the feature extraction block, a wavenet conversion model, a time-invariant synthesis filter, and a synthesized speech. In the training phase of the vocoder, the speech signal may be the initial input in the form of raw audio waveform (for example, speech). At the feature extraction block, convert the audio waveform into a mel-spectrogram to capture frequency and time information. At the time resolution adjustment block, adjust the time resolution of the extracted features. At the Mapping Technique Block, apply mapping techniques to correlate the features with the desired audio output characteristics. In the testing phase of the vocoder, the speech signal may be initial input for testing, similar to the training phase. the vocoder may transform the audio waveform into the mel-spectrogram for consistency with the training phase. At wavenet conversion model, uses the WaveNet architecture to generate audio samples based on the mel-spectrogram. At time-invariant synthesis filter, apply a synthesis filter to ensure that the generated audio maintains a natural and consistent sound. Finally, synthesized speech, the final output, which is the enhanced and improved speech audio waveform.
Figure 11 illustrates an example block diagram control method for the electronic apparatus.
In a figure 11, an electronic apparatus may include a memory, at least one processor comprising a processing circuit, wherein the at least one processor configured to obtain a content associated with a virtual reality environment (1110), wherein the content includes audio data and image data, obtain video frame information associated with the image data (1120), wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data, obtain audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed (1130), and obtain (or render) enhanced audio data by rendering the audio data based on the audio spatial characteristics (1140).
The content may refer to data associated with a virtual reality environment, such as audio data, image data, and metadata that collectively define a virtual experience. In addition, The content may include video frame information associated with the image data. The at least one processor may obtain the video frame information included in the content.
The video frame information may refer to data associated with one or more video frames, including a frame speed, timestamps, frame order, or other temporal parameters used for processing or rendering the video content.
The video frame information may be corresponding to the variable video frame information.
The video frame information may be represented as variable video frame information video frame metadata, frame timing information, video frame parameters, image frame attributes, or frame display information.
The video frame speed may refer to the display rate of video frames over time.
The video frame speed may corresponding to the speed of one or more variable image frames.
The video frame speed may be represented as a frame display rate, video playback speed, frame rate, visual playback rate, or image sequence speed.
The audio spatial characteristics may refer to properties of audio signals that define the perceived location, movement, and directionality of the sound source within a three-dimensional space.
The audio spatial characteristics may be represented as spatial sound properties, 3D audio attributes, sound spatial features or audio parameters.
The audio spatial characteristics may be corresponding to the at least one of localization and movement of sounds.
The audio spatial characteristics may be used for preserving one or more spatial audio cues.
The audio spatial characteristics may be used for ensuring that the audio processing model maintains a sense of immersion for a user during time-manipulated video frame information being experienced within the virtual reality environment.
The enhanced audio data may refer to audio data that has been processed or rendered to reflect spatial characteristics, improve clarity, or increase immersion within a virtual environment.
The enhanced audio data may refer to audio output data generated by applying spatial rendering or sound optimization techniques based on the virtual reality context.
The enhanced audio data may be represented as processed audio output, rendered sound data, spatialized audio stream, augmented audio content, or optimized audio signal.
The audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
The sound localization may refer to the process or ability to determine the position or direction of the sound source within a virtual space, based on auditory cues.
The sound localization may be represented as audio source positioning, sound source detection, acoustic localization, directional sound identification, or spatial audio pinpointing.
The sound movement may refer to the perceived change in position of the sound source within a virtual space.
The sound movement may refer to dynamic changes in the location or direction of the sound source.
The sound movement may be represented as audio source motion, dynamic sound positioning, moving sound effect, spatial audio transition, or sound trajectory.
The at least one processor may obtain an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtain the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
The audio processing model may be corresponding to the audio processing module (116).
The audio processing model may obtain a plurality of audio features based on the video frame information.
The at least one processor may obtain the audio spatial characteristics indicating the plurality of audio features based on the audio data and the video frame information.
The at least one processor may generate the audio spatial characteristics by using the audio processing model.
The plurality of audio features may be associated with the video frame speed.
The at least one processor may output the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
The at least one processor may generate the enhanced audio data, which has been processed based on spatial audio characteristics (e.g., direction, distance, movement of sound sources).
At least one processor may output enhanced audio data to a speaker of the electronic apparatus based on synchronization information, while image data may be displayed on a display of the electronic apparatus.
The enhanced audio data may reflect spatial characteristics of sound, such as direction, distance, and movement, and the image data may be associated with a virtual reality environment.
The at least one processor may also receive synchronization information, such as timestamps or frame rate data, that indicates how the audio and image data are temporally related.
The at least one processor may control output timing of the enhanced audio data to align with the image data. For example, when a sound source appears in a specific frame, the corresponding sound may be output at the exact time the frame is displayed.
The enhanced audio data may provide the realism and immersion of the user experience in a virtual environment.
The at least one processor may identify a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtain the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
The sound spatial position may refer to a location of a sound source in three-dimensional space within a virtual reality environment, identified based on the video frame information and the audio spatial characteristics.
The sound spatial position is a specific data output representing the physical or virtual coordinates of a sound source.
The Audio spatial characteristics influence perception, while the sound spatial position provides a basis for rendering accurate sound placement.
The audio spatial characteristics may represent a set of extracted or derived parameters from audio data, such as directionality vectors, interaural time differences (ITD), interaural level differences (ILD), or head-related transfer function (HRTF) profiles. These parameters are used to model how sound should be rendered to simulate a spatial environment.
The sound spatial position may refer to a concrete computation result (typically a set of three-dimensional coordinates (x, y, z)) that defines the specific location of a sound source in the virtual environment. The at least one processor may obtain the sound spatial position based on both the audio spatial characteristics and additional contextual data such as video frame information or user orientation.
The at least one processor may use audio spatial characteristics as input features for spatial modeling. The at least one processor may use the sound spatial position for rendering spatial audio aligned with visual content.
The sound spatial position may indicate a location of the sound source in three-dimensional space relative to a user's position.
The at least one processor may update the sound spatial position in real-time based on a change of the video frame speed.
The at least one processor may identify the change of the video frame speed. Based on the change of the video frame speed, the at least one processor may update the sound spatial position based on the changed video frame speed.
The at least one processor may obtain the audio spatial characteristics based on one or more spatial audio cues in the audio data, wherein the one or more spatial audio cues may include at least one of increasing or decreasing speed associated with a sound source.
The at least one processor may obtain the audio spatial characteristics based on at least one visual cues in the image data, wherein the at least one visual cues may include information related with an object corresponding to a sound source.
The at least one visual cues within a current scene may provide information about the virtual reality environment and help the user make sense of spatial relationships, motion, and positioning of objects and sounds.
The at least one processor may identify motion sickness degree of a user based on the video frame information associated with the virtual reality environment, and obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the motion sickness degree.
The at least one processor may identify(or predict) the motion sickness of the user based on the image data. The at least one processor may analyze frames of the image data to predict a user's degree of motion sickness based on visual motion patterns or frame characteristics.
The electronic apparatus may store data in advance to determine whether the pattern is likely to cause motion sickness. The electronic apparatus may analyze a pattern representing screen transitions based on image frames and the stored data.
The at least one processor may obtain movement degree of a user. The at least one processor may obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the movement degree.
The at least one processor may obtain a sensing data for user’s motion. The electronic apparatus may include a sensor configured to sense a user’s motion. For example, the sensor is IMU (Inertial Measurement Unit), accelerometer sensor, gyroscope sensor or image sensor.
The at least one processor may obtain the movement degree of the user based on the sensing data.
In an embodiment, an method for controlling an electronic apparatus, the method comprising: obtaining a content associated with a virtual reality environment, wherein the content may include audio data and image data, obtaining video frame information associated with the image data, wherein the video frame information may include a video frame speed for displaying at least one video frame included in the image data, obtaining audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and obtaining enhanced audio data by rendering the audio data based on the audio spatial characteristics.
The audio spatial characteristics may include at least one of a sound localization and a sound movement associated with a sound source.
The obtaining the audio spatial characteristics may include obtaining an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and obtaining the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
The method may include outputting the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
The obtaining enhanced audio data may include identifying a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and obtaining the enhanced audio data based on the sound spatial position, wherein the sound spatial position may include three-dimensional coordinates associated with a sound source.
The software may comprise an ordered listing of executable instructions for implementing logical functions, and may be embodied in any "processor-readable medium" for use by or in connection with an instruction execution system, apparatus, or device, such as a single or multiple-core processor or processor-containing system.
The blocks or operations of a method or algorithm and functions described in connection with embodiments disclosed herein may be embodied directly in hardware, in a software module executed by a processor, or in a combination of the two. If implemented in software, the functions may be stored on or transmitted over as one or more instructions or code on a tangible, non-transitory computer-readable medium (e.g., the memory 302 and/or the memory 302). A software module may reside in Random Access Memory (RAM), flash memory, Read Only Memory (ROM), Electrically Programmable ROM (EPROM), Electrically Erasable Programmable ROM (EEPROM), registers, hard disk, a removable disk, a CD ROM, or any other form of storage medium known in the art.
Embodiments of the present disclosure provide a system that minimizes distortion and unnatural sounds often associated with time-warped video playback. The system preserves the immersive and believable spatial audio experience crucial for VR or 360° video. Further, the system reduces motion sickness by eliminating audio-visual discrepancies that can contribute to discomfort. In AR or VR, a time-manipulated variable video is played to reduce motion sickness and the audio is not distorted. While sickness is managed through a low-latency image rendering to the latest viewport, appropriate audio is not generated in this manner to preserve the immersive experience of the AR or VR. The system tracks sound localization and movement in response to video speed changes. Further, the system adapts audio playback speed, resamples audio signals, and applies spatial audio effects for seamless immersion. Furthermore, the system addresses the challenge of maintaining immersive audio during time-manipulated VR and mitigates motion sickness.
Embodiments disclosed herein may be implemented through at least one software program running on at least one hardware device and performing network management functions to control the elements. The elements may be at least one of a hardware device, or a combination of hardware device and software module.
Unless otherwise defined, all technical and scientific terms used herein have the same meaning as, or a similar meaning to, that commonly understood by one ordinary skilled in the art to which the inventive concepts belong. The system, methods, and examples provided herein are illustrative only and not intended to be limiting.
While specific language has been used to describe the present subject matter, any limitations arising on account thereto, are not intended. As would be apparent to a person in the art, various working modifications may be made to the method to implement the inventive concepts as taught herein. The drawings and the forgoing description give examples of embodiments. Those skilled in the art will appreciate that one or more of the described elements may well be combined into a single functional element. Alternatively, certain elements may be split into multiple functional elements. Elements from one example may be added to another example.
Embodiments disclosed herein may be implemented using at least one hardware device and performing network management functions to control the elements.
The foregoing description of the specific examples will so fully reveal the general nature of embodiments herein that others may, by applying current knowledge, readily modify and/or adapt for various applications such specific examples without departing from the generic concept, and, therefore, such adaptations and modifications should and are intended to be comprehended within the meaning and range of equivalents of embodiments. It is to be understood that the phraseology or terminology employed herein is for the purpose of description and not of limitation. Therefore, while embodiments herein have been described in terms of examples, those skilled in the art will recognize that the examples herein may be practiced with modification within the scope of embodiments as described herein.

Claims (15)

  1. An electronic apparatus comprising:
    at least one processor comprising a processing circuit;
    memory storing instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    obtain a content associated with a virtual reality environment, wherein the content includes audio data and image data,
    obtain video frame information associated with the image data, wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data,
    obtain audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and
    obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics.
  2. The electronic apparatus of claim 1, wherein the audio spatial characteristics includes at least one of a sound localization and a sound movement associated with a sound source.
  3. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    obtain an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and
    obtain the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
  4. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    output the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  5. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    identify a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and
    obtain the enhanced audio data based on the sound spatial position,
    wherein the sound spatial position includes three-dimensional coordinates associated with a sound source.
  6. The electronic apparatus of claim 5, wherein the sound spatial position indicates a location of the sound source in three-dimensional space relative to a user's position.
  7. The electronic apparatus of claim 5, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    update the sound spatial position in real-time based on a change of the video frame speed.
  8. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    obtain the audio spatial characteristics based on one or more spatial audio cues in the audio data,
    wherein the one or more spatial audio cues includes at least one of increasing or decreasing speed associated with a sound source.
  9. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    obtain the audio spatial characteristics based on at least one visual cues in the image data,
    wherein the at least one visual cues includes information related with an object corresponding to a sound source.
  10. The electronic apparatus of claim 1, wherein the instructions that, when executed by the at least one processor individually or collectively , cause the electronic device to:
    identify motion sickness degree of a user based on the video frame information associated with the virtual reality environment, and
    obtain enhanced audio data by rendering the audio data based on the audio spatial characteristics and the motion sickness degree.
  11. A method for controlling an electronic apparatus, the method comprising:
    obtaining a content associated with a virtual reality environment, wherein the content includes audio data and image data,
    obtaining video frame information associated with the image data, wherein the video frame information includes a video frame speed for displaying at least one video frame included in the image data,
    obtaining audio spatial characteristics in the virtual reality environment based on the audio data and the video frame speed, and
    obtaining enhanced audio data by rendering the audio data based on the audio spatial characteristics.
  12. The method of claim 11, wherein the audio spatial characteristics includes at least one of a sound localization and a sound movement associated with a sound source.
  13. The method of claim 11, wherein the obtaining the audio spatial characteristics comprises:
    obtaining an audio processing model to obtain the audio spatial characteristics in the virtual reality environment, and
    obtaining the audio spatial characteristics by inputting the audio data and the video frame speed into the audio processing model.
  14. The method of claim 11, wherein the method further comprising:
    outputting the enhanced audio data a speaker of the electronic apparatus based on synchronization information while displaying the image data on a display of the electronic apparatus.
  15. The method of claim 11, wherein the obtaining enhanced audio data comprises:
    identifying a sound spatial position within the virtual reality environment the based on the video frame information and the audio spatial characteristics, and
    obtaining the enhanced audio data based on the sound spatial position,
    wherein the sound spatial position includes three-dimensional coordinates associated with a sound source.
PCT/KR2025/007342 2024-06-25 2025-05-29 System and method of enhanced audio rendering for an immersive virtual reality experience Pending WO2026005309A1 (en)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
IN202411048655 2024-06-25
IN202411048655 2024-06-25

Publications (1)

Publication Number Publication Date
WO2026005309A1 true WO2026005309A1 (en) 2026-01-02

Family

ID=98222340

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/KR2025/007342 Pending WO2026005309A1 (en) 2024-06-25 2025-05-29 System and method of enhanced audio rendering for an immersive virtual reality experience

Country Status (1)

Country Link
WO (1) WO2026005309A1 (en)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120134543A1 (en) * 2010-11-30 2012-05-31 Fedorovskaya Elena A Method of identifying motion sickness
US20180307305A1 (en) * 2017-04-24 2018-10-25 Intel Corporation Compensating for High Head Movement in Head-Mounted Displays
US20220319014A1 (en) * 2021-04-05 2022-10-06 Facebook Technologies, Llc Systems and methods for dynamic image processing and segmentation
CN115187899A (en) * 2022-07-04 2022-10-14 京东科技信息技术有限公司 Audio and video synchronization judging method and device, electronic equipment and storage medium
US20240107113A1 (en) * 2022-09-22 2024-03-28 Apple Inc. Parameter Selection for Media Playback

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20120134543A1 (en) * 2010-11-30 2012-05-31 Fedorovskaya Elena A Method of identifying motion sickness
US20180307305A1 (en) * 2017-04-24 2018-10-25 Intel Corporation Compensating for High Head Movement in Head-Mounted Displays
US20220319014A1 (en) * 2021-04-05 2022-10-06 Facebook Technologies, Llc Systems and methods for dynamic image processing and segmentation
CN115187899A (en) * 2022-07-04 2022-10-14 京东科技信息技术有限公司 Audio and video synchronization judging method and device, electronic equipment and storage medium
US20240107113A1 (en) * 2022-09-22 2024-03-28 Apple Inc. Parameter Selection for Media Playback

Similar Documents

Publication Publication Date Title
Gan et al. Music gesture for visual sound separation
Gan et al. Foley music: Learning to generate music from videos
CN113299312B (en) Image generation method, device, equipment and storage medium
SG11202108498RA (en) Method and device for generating video, electronic equipment, and computer storage medium
WO2022110354A1 (en) Video translation method, system and device, and storage medium
WO2022260432A1 (en) Method and system for generating composite speech by using style tag expressed in natural language
WO2022029044A1 (en) Method and electronic device
WO2020017798A1 (en) A method and system for musical synthesis using hand-drawn patterns/text on digital and non-digital surfaces
CN102087704A (en) Information processing apparatus, information processing method, and program
Mo et al. A unified audio-visual learning framework for localization, separation, and recognition
WO2023101377A1 (en) Method and apparatus for performing speaker diarization based on language identification
WO2022071959A1 (en) Audio-visual hearing aid
Chen et al. Sound localization by self-supervised time delay estimation
Li et al. Audiovisual source association for string ensembles through multi-modal vibrato analysis
Montesinos et al. Solos: A dataset for audio-visual music analysis
US20130218570A1 (en) Apparatus and method for correcting speech, and non-transitory computer readable medium thereof
WO2022059869A1 (en) Device and method for enhancing sound quality of video
CA3184814A1 (en) A system (variants) for providing a harmonious combination of video files and audio files and a related method
Sudo et al. Environmental sound segmentation utilizing Mask U-Net
WO2024205147A1 (en) Method and server for providing media content
WO2026005309A1 (en) System and method of enhanced audio rendering for an immersive virtual reality experience
Sudo et al. Multi-channel environmental sound segmentation
Li et al. Online audio-visual source association for chamber music performances
Sarasúa Context-aware gesture recognition in classical music conducting
US20240080566A1 (en) System and method for camera handling in live environments

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25827172

Country of ref document: EP

Kind code of ref document: A1