HK40112923A - Soundscape augmentation system and method of forming the same - Google Patents

Soundscape augmentation system and method of forming the same Download PDF

Info

Publication number
HK40112923A
HK40112923A HK62024100927.8A HK62024100927A HK40112923A HK 40112923 A HK40112923 A HK 40112923A HK 62024100927 A HK62024100927 A HK 62024100927A HK 40112923 A HK40112923 A HK 40112923A
Authority
HK
Hong Kong
Prior art keywords
soundscape
data
masking
various embodiments
enhancement system
Prior art date
Application number
HK62024100927.8A
Other languages
Chinese (zh)
Inventor
黄文睿
黄株绮
林滨
王镇廷
黄智明
颜允圣
Original Assignee
南洋理工大学
建屋发展局
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by 南洋理工大学, 建屋发展局 filed Critical 南洋理工大学
Publication of HK40112923A publication Critical patent/HK40112923A/en

Links

Description

声景增强系统及其形成方法Soundscape enhancement system and its formation method

相关申请的交叉引用Cross-references to related applications

本申请要求于2022年4月27日提交的新加坡申请号10202204451S的优先权的权益,其内容在此通过援引整体并入本文,以用于所有目的。This application claims the benefit of priority to Singapore application No. 10202204451S, filed on 27 April 2022, the contents of which are incorporated herein by reference in their entirety for all purposes.

技术领域Technical Field

本申请的实施例涉及声景增强系统。本申请的实施例涉及声景增强系统的形成方法。Embodiments of this application relate to soundscape enhancement systems. Embodiments of this application also relate to methods for forming soundscape enhancement systems.

背景技术Background Technology

世界卫生组织(World Health Organization,WHO)已经将噪声暴露列为排在空气污染之后的第二大环境污染。噪声暴露也被比作二手烟。采取行动以减轻噪声暴露的迫切呼吁源于噪声对健康的影响的、越来越多的证据,噪声对健康的影响是例如缺血性心脏病发病率的上升风险、烦躁和睡眠障碍。世界卫生组织还强调环境噪声暴露对心理健康和幸福感(well-being)的影响,其中新证据表明对抑郁和焦虑指标的不利影响。The World Health Organization (WHO) has listed noise exposure as the second leading cause of environmental pollution after air pollution. Noise exposure has also been compared to secondhand smoke. The urgent call to take action to mitigate noise exposure stems from mounting evidence of the health effects of noise, such as an increased risk of ischemic heart disease, irritability, and sleep disturbances. The WHO also emphasizes the impact of environmental noise exposure on mental health and well-being, with new evidence highlighting its adverse effects on indicators of depression and anxiety.

从噪声管理到声景管理的范式转变(paradigm shift)正在出现。ISO12913-1将“声景”定义为“所感知的或所体验的声学环境和/或被一个人或多个人在情境中所理解的声学环境”。与噪声管理相反,声景管理框架将声音视为资源而不是废物;声景管理框架专注于偏爱的声音而不是令人不适的声音;并且声景管理框架以期望出现的声音进行管理,以掩蔽不期望出现的声音,并且管理以减少不期望出现的声音,以代替仅降低声级(soundlevels)。因此,有基于“掩蔽”声音的增强或在声学环境中引入“掩蔽”声音的声景干预技术,以改善声学舒适度的整体感知。通常,该干预涉及自然声音的增强(例如通过扬声器),并且已在户外休闲场所和例如疗养院的室内进行了试验,以提高声学舒适度。利用期望出现的声音的声景增强与声音掩蔽系统相似,声音掩蔽系统通常在办公室环境中使用,以减少干扰。声音掩蔽系统与基于声景的增强系统之间的关键区别在于掩蔽是基于客观指标,然而声景增强则主要基于以情境为基础的主观指标。情境“包括人与活动和地点之间的、在空间和时间上的相互关系”,并且在ISO12913-1中通过示例进一步详细说明。还值得注意的是,声景不仅限于现实世界的环境,还可以包括虚拟环境,并且甚至可以包括记忆中的回忆。A paradigm shift from noise management to soundscape management is emerging. ISO 12913-1 defines "soundscape" as "the perceived or experienced acoustic environment and/or the acoustic environment as understood by one or more people in a context." In contrast to noise management, the soundscape management framework treats sound as a resource, not waste; it focuses on preferred sounds rather than unpleasant ones; and it manages unwanted sounds by masking unwanted ones, and by reducing unwanted sounds rather than simply lowering sound levels. Therefore, soundscape intervention techniques exist that enhance or introduce "masking" sounds into the acoustic environment to improve the overall perceived acoustic comfort. Typically, this intervention involves enhancing natural sounds (e.g., through loudspeakers) and has been tested in outdoor recreational spaces and indoor spaces such as nursing homes to improve acoustic comfort. Soundscape enhancement utilizing desired sounds is similar to sound masking systems, which are commonly used in office environments to reduce disturbances. The key difference between sound masking systems and soundscape-based enhancement systems is that masking is based on objective metrics, while soundscape enhancement is primarily based on context-based subjective metrics. Context "includes the spatial and temporal interrelationships between people and activities and places," and is further described in detail by example in ISO 12913-1. It is also worth noting that soundscapes are not limited to real-world environments; they can include virtual environments and even memories.

发明内容Summary of the Invention

本申请的实施例可以提供声景增强系统。声景增强系统可以包括数据获取系统,数据获取系统被设置为提供周围声景数据。声景增强系统还可以包括数据库,数据库包括多个掩蔽设置。声景增强系统还可以包括感知属性预测器,感知属性预测器耦接于数据获取系统和数据库,感知属性预测器被设置为基于周围声景数据,为多个掩蔽设置中的每个掩蔽设置,生成在一个或多个预定义感知属性指标上代表感知的预测。声景增强系统可以附加地包括掩蔽设置排序系统,掩蔽设置排序系统被设置为基于感知属性预测器生成的预测,确定一个或多个最佳掩蔽设置。声景增强系统还可以包括播放系统,播放系统被设置为播放或再现一个或多个最佳掩蔽设置。Embodiments of this application may provide a soundscape enhancement system. The soundscape enhancement system may include a data acquisition system configured to provide ambient soundscape data. The soundscape enhancement system may also include a database comprising multiple masking settings. The soundscape enhancement system may further include a perception attribute predictor coupled to the data acquisition system and the database, configured to generate, based on the ambient soundscape data, a prediction representing perception on one or more predefined perception attribute metrics for each of the multiple masking settings. The soundscape enhancement system may additionally include a masking setting ranking system configured to determine one or more optimal masking settings based on the predictions generated by the perception attribute predictor. The soundscape enhancement system may also include a playback system configured to play back or reproduce one or more optimal masking settings.

本申请的实施例可以涉及声景增强系统的形成方法。该方法可以包括提供数据获取系统,数据获取系统被设置为提供周围声景数据。该方法还可以包括提供数据库,数据库包括多个掩蔽设置。该方法可以进一步包括将感知属性预测器耦接于数据获取系统和数据库,感知属性预测器被设置为基于周围声景数据,为多个掩蔽设置中的每个掩蔽设置,生成在一个或多个预定义感知属性指标上代表感知的预测。该方法可以附加地包括提供掩蔽设置排序系统,掩蔽设置排序系统被设置为基于感知属性预测器生成的预测,确定一个或多个最佳掩蔽设置。该方法可以进一步包括提供播放系统,播放系统被设置为播放或再现一个或多个最佳掩蔽设置。Embodiments of this application may relate to a method for forming a soundscape enhancement system. The method may include providing a data acquisition system configured to provide ambient soundscape data. The method may also include providing a database comprising a plurality of masking settings. The method may further include coupling a perceptual attribute predictor to the data acquisition system and the database, the perceptual attribute predictor being configured to generate, based on the ambient soundscape data, a prediction representing perception on one or more predefined perceptual attribute metrics for each of the plurality of masking settings. The method may additionally include providing a masking setting ranking system configured to determine one or more optimal masking settings based on the predictions generated by the perceptual attribute predictor. The method may further include providing a playback system configured to play back or reproduce one or more optimal masking settings.

附图说明Attached Figure Description

在附图中,相似的参考符号在不同视图中通常指相同的部件。附图不一定按比率绘制,而是实质上着重于说明各个实施例的原理。在以下描述中,将参考以下附图描述本发明的各个实施例。In the accompanying drawings, similar reference numerals generally refer to the same parts in different views. The drawings are not necessarily drawn to scale, but rather focus on illustrating the principles of various embodiments. In the following description, various embodiments of the invention will be described with reference to the following drawings.

图1示出根据各个实施例的声景增强系统的示意图。Figure 1 shows a schematic diagram of a soundscape enhancement system according to various embodiments.

图2示出根据各个实施例的、声景增强系统的形成方法的示意图。Figure 2 shows a schematic diagram of a method for forming a soundscape enhancement system according to various embodiments.

图3示出根据各个实施例的声景增强系统的示意图。Figure 3 shows a schematic diagram of a soundscape enhancement system according to various embodiments.

图4示出ISO12913-3提出的情感质量属性的二维圆周八分模型,根据各个实施例的感知属性预测器的预测可以基于该二维圆周八分模型。Figure 4 illustrates a two-dimensional circular octet model of the emotional quality attributes proposed in ISO 12913-3. The predictions of the perceptual attribute predictors according to various embodiments can be based on this two-dimensional circular octet model.

图5是根据各个实施例的感知属性预测器的示意图。Figure 5 is a schematic diagram of a perceptual attribute predictor according to various embodiments.

图6是根据各个实施例的感知属性预测器的示意图。Figure 6 is a schematic diagram of a perceptual attribute predictor according to various embodiments.

图7是根据各个实施例的感知属性预测器的示意图。Figure 7 is a schematic diagram of a perceptual attribute predictor according to various embodiments.

图8示出根据各个实施例的自动掩蔽选择系统(automatic masker selectionsystem,AMSS)的训练和推理方案。Figure 8 illustrates the training and inference schemes for the automatic mask selection system (AMSS) according to various embodiments.

图9示出(a)根据各个实施例使用的基本卷积循环神经网络(convolutionalrecurrent neural network,CRNN)架构;(b)根据各个实施例的、(a)中的特征映射块的、一个可能的实施;以及(c)根据各个其他实施例的、(a)中的特征映射块的、其他可能的实施。Figure 9 illustrates (a) the basic convolutional recurrent neural network (CRNN) architecture used according to various embodiments; (b) a possible implementation of the feature map block in (a) according to various embodiments; and (c) other possible implementations of the feature map block in (a) according to various other embodiments.

图10是示出根据各个实施例的、对每个设置测试的10次运行(上:交叉验证集,下:测试集)中的概率感知属性预测器(probabilistic perceptual attribute predictor,PPAP)的均折(mean fold)均方误差(mean squared errors,MSEs)(±标准差)的表格。Figure 10 is a table showing the mean fold mean squared errors (MSEs) (± standard deviation) of the probabilistic perceptual attribute predictor (PPAP) in 10 runs (top: cross-validation set, bottom: test set) for each setting according to various embodiments.

图11示出使用测试集上的50个模型(点积变量)进行的掩蔽选择,其中具有根据各个实施例的、朴素最大μk值选择方案(左)以及根据各个实施例的随机采样方案,以激励掩蔽器探索(右)。Figure 11 illustrates mask selection using 50 models (dot product variables) on the test set, with a naive maximum μk value selection scheme (left) according to various embodiments and a random sampling scheme (right) according to various embodiments to incentivize masker exploration.

图12A示出根据各个实施例的概率感知属性预测器(probabilistic perceptualattribute predictor,PPAP)的示意图。Figure 12A shows a schematic diagram of a probabilistic perceptual attribute predictor (PPAP) according to various embodiments.

图12B示出各个实施例中根据计算式(8)、(9)和(10)的三种特征增强方法。Figure 12B illustrates three feature enhancement methods based on calculation formulas (8), (9), and (10) in various embodiments.

图13示出根据各个实施例的、减少总运行时间的算法。Figure 13 illustrates algorithms for reducing total runtime according to various embodiments.

图14示出(上图)均方误差(mean squared error,MSE)作为注意力块类型的函数的图示,该图示出根据各个实施例的、每个注意力块类型和特征增强方法的愉悦度预测的验证MSE的小提琴图示;以及(下图)平均绝对误差(mean absolute error,MAE)作为注意力块类型和特征增强方法的函数的图示,该图示出根据各个实施例的、愉悦度预测的验证MAE的小提琴图示。Figure 14 shows (top) a diagram illustrating mean squared error (MSE) as a function of attention block type, illustrating a violin diagram of the validation MSE for pleasure prediction according to various embodiments for each attention block type and feature enhancement method; and (bottom) a diagram illustrating mean absolute error (MAE) as a function of attention block type and feature enhancement method, illustrating a violin diagram of the validation MAE for pleasure prediction according to various embodiments.

图15A是根据各个实施例的、ISO愉悦度作为对数增益(log gain)的函数的图示,该图示出使用加性注意力(additive attention,AA)和CONV增强(种子5、验证折2)的、为掩蔽器类别“水”进行的增益插值(gain interpolation)。Figure 15A is a diagram illustrating ISO pleasure as a function of log gain according to various embodiments. The diagram shows gain interpolation for the masker category "water" using additive attention (AA) and CONV enhancements (seed 5, verification fold 2).

图15B是根据各个实施例的、ISO愉悦度作为对数增益的函数的图示,该图示出使用具有加性注意力(AA)和CONV增强(种子5,验证折2)的、为掩蔽器类别“交通”进行的增益插值。Figure 15B is a diagram illustrating ISO pleasure as a function of logarithmic gain according to various embodiments, showing gain interpolation for the masker category "traffic" using additive attention (AA) and CONV enhancement (seed 5, verification fold 2).

图15C是根据各个实施例的、ISO愉悦度作为对数增益的函数的图示,该图示出使用具有加性注意力(AA)和CONV增强(种子5、验证折2)的、为掩蔽器类别“鸟”进行的增益插值。Figure 15C is a diagram illustrating ISO pleasure as a function of logarithmic gain according to various embodiments, showing gain interpolation for the masker category "bird" using additive attention (AA) and CONV enhancements (seed 5, verification fold 2).

图15D是根据各个实施例的、ISO愉悦度作为对数增益的函数的图示,该图示出使用加性注意力(AA)和CONV增强(种子5、验证折2)的、为掩蔽器类别“施工”进行的增益插值。Figure 15D is a diagram of ISO pleasure as a function of logarithmic gain according to various embodiments, illustrating gain interpolation for the masker category "Construction" using additive attention (AA) and CONV enhancements (seed 5, verification fold 2).

图15E是根据各个实施例的、ISO愉悦度作为对数增益的函数的图示,该图示出使用加性注意力(AA)和CONV增强(种子5、验证折2)的、为掩蔽器类别“静音”进行的增益插值。Figure 15E is a diagram of ISO pleasure as a function of logarithmic gain according to various embodiments, illustrating gain interpolation for the masker category "Mute" using additive attention (AA) and CONV enhancement (seed 5, verification fold 2).

具体实施方式Detailed Implementation

以下详细描述参考附图,附图以说明的方式示出可以实施本发明的具体细节和实施例。这些实施例被足够详细地描述,以使本领域技术人员能够实施本发明。可以使用其他实施例,并且可以在不脱离本发明范围的情况下,进行结构方面、逻辑方面和电方面的改变。各个实施例不一定是相互排斥的,因为一些实施例可以与一个或多个其他实施例组合以形成新的实施例。The following detailed description refers to the accompanying drawings, which illustrate specific details and embodiments in which the invention can be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the invention. Other embodiments may be used, and structural, logical, and electrical changes may be made without departing from the scope of the invention. The various embodiments are not necessarily mutually exclusive, as some embodiments may be combined with one or more other embodiments to form new embodiments.

在一个实施例的上下文中描述的特征可以相应地适用于其他实施例中相同或相似的特征。在一个实施例的上下文中描述的特征可以相应地适用于其他实施例,即使在这些其他实施例中没有明确描述。此外,在一个实施例的上下文中针对一个特征描述的添加和/或组合和/或替代可以相应地适用于其他实施例中相同或相似的特征。Features described in the context of one embodiment may be adapted accordingly to the same or similar features in other embodiments. Features described in the context of one embodiment may be adapted accordingly to other embodiments, even if not explicitly described in those other embodiments. Furthermore, additions and/or combinations and/or substitutions described for a feature in the context of one embodiment may be adapted accordingly to the same or similar features in other embodiments.

在各个实施例的上下文中,关于特征或元件使用的冠词“一”、“一个”和“该”包括对一个或多个特征或元件的引用。In the context of the various embodiments, the articles “a,” “an,” and “the” used with respect to features or elements include references to one or more features or elements.

在各个实施例的上下文中,应用于数值的术语“约”或“大约”涵盖确切值和合理差值,合理差值在例如指定值的10%以内。In the context of the various embodiments, the term “about” or “approximately” applied to numerical values covers both exact values and reasonable differences, where a reasonable difference is, for example, within 10% of a specified value.

如本文所用,术语“和/或”包括所列相关项目中的一个或多个项目的任意组合和全部组合。As used herein, the term “and/or” includes any and all combinations of one or more of the listed related items.

“包括”指包括但不限于“包括”一词后面的任何内容。因此,术语“包括”的使用表示所列元件是必需的或强制性的,但是其他元件是可选的,并且其他元件可以存在或不存在。"Including" means, but is not limited to, anything that follows the word "including". Therefore, the use of the term "including" indicates that the listed elements are necessary or mandatory, but other elements are optional and may or may not be present.

“由……组成”指包括且限于短语“由……组成”之间的任何内容。因此,短语“由……组成”表示所列元件是必需的或强制性的,并且不存在其他元件。"Composed of" means anything included in and limited to the phrase "composed of". Therefore, the phrase "composed of" indicates that the listed elements are necessary or mandatory, and that no other elements exist.

针对一个声景增强系统的上下文中描述的实施例对于其他声景增强系统同样有效。类似地,在方法的上下文中描述的实施例对于声景增强系统同样有效,反之亦然。The embodiments described in the context of one soundscape enhancement system are equally effective for other soundscape enhancement systems. Similarly, the embodiments described in the context of a method are equally effective for soundscape enhancement systems, and vice versa.

目前最先进的增强或掩蔽干预涉及通过扬声器的声音静态播放,其中必须手动选择声级和音轨。各个实施例可以利用首创的人工智能模型,以自动选择音轨,该音轨在最合适的声级下产生任意主观指标(例如感知响度、ISO12913-3的圆周八分指标(CircumplexOctant Scale)-愉悦度(pleasantness)、烦扰度(annoyance)、宁静度(tranquillity)、平静度(calmness)、活跃度(vibrancy)、事件性(eventfulness)等)上的最高(或最低)值。人工智能(artificial intelligence,AI)模型是由南洋理工大学(Nanyang TechnologicalUniversity,NTU)研究团队在基于ISO12913-2标准收集的新加坡当地人群的人类主观反应的大数据集上进行训练的。ISO12913系列标准代表声音环境管理的范式转变,并且详细描述基于感知的方法以全面评估和分析声音环境。此外,人工智能模型还可以在不按照ISO12913-2的要求进行繁琐且耗力的调查或问卷的情况下,兼作自动评估声音环境感知质量的预测工具。The most advanced enhancement or masking interventions currently involve static sound playback through loudspeakers, where sound levels and tracks must be manually selected. Various embodiments can utilize a pioneering artificial intelligence model to automatically select a track that produces the highest (or lowest) value on any subjective metric (e.g., perceived loudness, the ISO 12913-3 Circumplex Octant Scale – pleasantness, annoyance, tranquility, calmness, vibrancy, eventfulness, etc.) at the most suitable sound level. The artificial intelligence (AI) model was trained by a research team at Nanyang Technological University (NTU) on a large dataset of human subjective responses from the local population in Singapore, collected based on the ISO 12913-2 standard. The ISO 12913 series of standards represents a paradigm shift in sound environment management and details a perception-based approach for the comprehensive assessment and analysis of sound environments. In addition, artificial intelligence models can also serve as predictive tools for automatically assessing the perceived quality of the sound environment without having to conduct tedious and laborious surveys or questionnaires as required by ISO 12913-2.

图1是根据各个实施例的声景增强系统的示意图。声景增强系统可以包括数据获取系统102,数据获取系统102被设置为提供周围声景数据。声景增强系统还可以包括数据库104,数据库104包括多个掩蔽设置。声景增强系统可以进一步包括感知属性预测器106,感知属性预测器106耦接于数据获取系统102和数据库104,感知属性预测器106被设置为基于周围声景数据,为多个掩蔽设置中的每个掩蔽设置,生成在一个或多个预定义感知属性指标上代表感知的预测。声景增强系统可以附加地包括掩蔽设置排序系统108,掩蔽设置排序系统108被设置为基于感知属性预测器106生成的预测,确定一个或多个最佳掩蔽设置。声景增强系统还可以包括播放系统110,播放系统110被设置为播放或再现一个或多个最佳掩蔽设置。Figure 1 is a schematic diagram of a soundscape enhancement system according to various embodiments. The soundscape enhancement system may include a data acquisition system 102 configured to provide ambient soundscape data. The soundscape enhancement system may also include a database 104 including a plurality of masking settings. The soundscape enhancement system may further include a perception attribute predictor 106 coupled to the data acquisition system 102 and the database 104, configured to generate a prediction representing perception on one or more predefined perception attribute metrics for each of the plurality of masking settings based on the ambient soundscape data. The soundscape enhancement system may additionally include a masking setting ranking system 108 configured to determine one or more optimal masking settings based on the predictions generated by the perception attribute predictor 106. The soundscape enhancement system may also include a playback system 110 configured to play or reproduce one or more optimal masking settings.

换言之,声景增强系统可以包括感知属性预测器106,感知属性预测器106可以连接于数据获取系统102和数据库104。声景增强系统还可以包括连接于感知属性预测器106的掩蔽设置排序系统108和连接于掩蔽设置排序系统108的播放系统110。In other words, the soundscape enhancement system may include a perceptual attribute predictor 106, which may be connected to the data acquisition system 102 and the database 104. The soundscape enhancement system may also include a masking setting sorting system 108 connected to the perceptual attribute predictor 106 and a playback system 110 connected to the masking setting sorting system 108.

为避免疑问,图1旨在说明根据各个实施例的声景增强系统的一些特征,并且不旨在限制例如各个组件的尺寸、形状、取向、设置等。To avoid any doubt, Figure 1 is intended to illustrate some features of the soundscape enhancement system according to various embodiments, and is not intended to limit, for example, the size, shape, orientation, or arrangement of the various components.

在各个实施例中,数据获取系统102还可以被设置为提供群体统计或情境数据。感知属性预测器106还可以被设置为基于群体统计或情境数据生成预测。群体统计或情境数据可以由数据获取系统102直接接收、测量或感测,或者从外部源接收,外部源是例如便携式健康监测设备或存储设备等设备,如下文更详细描述的。In various embodiments, the data acquisition system 102 may also be configured to provide population statistics or contextual data. The perceived attribute predictor 106 may also be configured to generate predictions based on the population statistics or contextual data. The population statistics or contextual data may be received, measured, or sensed directly by the data acquisition system 102, or received from an external source, such as a portable health monitoring device or a storage device, as described in more detail below.

群体统计或情境数据可以包括与环境参数相关的数据、一名或多名听众的群体统计数据、一名或多名听众的心理数据、一名或多名听众的生理数据或其任意组合。Group statistics or contextual data may include data related to environmental parameters, group statistics of one or more listeners, psychological data of one or more listeners, physiological data of one or more listeners, or any combination thereof.

就此而言,一名或多名听众可以包括真实的听众和/或假想的听众。In this regard, one or more listeners may include real listeners and/or imaginary listeners.

在各个实施例中,与环境参数相关的环境数据可以是或者可以包括与位置相关的数据、与视觉环境相关的数据、附近的人数、空气温度、湿度、风速或其任意组合。环境参数可以是非声学环境参数。例如,数据获取系统102可以包括或者可以耦接于全球定位系统(Global Positioning System,GPS)传感器,以确定与位置相关的数据。数据获取系统102可以包括或者可以耦接于摄像头(camera),以确定与附近人数相关的数据和/或从环境中获得视觉信息(例如,确定特定类型的家具、宠物等的存在,确定人的移动,确定存在的绿化物或水的百分比等)。数据获取系统102可以包括或者可以耦接于温度计,以确定与空气温度相关的数据。数据获取系统102可以包括或者可以耦接于湿度计,以确定与湿度相关的数据。数据获取系统102可以包括或者可以耦接于风速计,以确定与风速相关的数据。在该上下文中,术语“耦接”可以指直接连接或间接连接。连接可以通过例如有线方式或无线方式建立。In various embodiments, environmental data related to environmental parameters may or may include location-related data, visual environment-related data, the number of people nearby, air temperature, humidity, wind speed, or any combination thereof. Environmental parameters may be non-acoustic environmental parameters. For example, data acquisition system 102 may include or be coupled to a Global Positioning System (GPS) sensor to determine location-related data. Data acquisition system 102 may include or be coupled to a camera to determine data related to the number of people nearby and/or to obtain visual information from the environment (e.g., determining the presence of specific types of furniture, pets, etc., determining human movement, determining the percentage of greenery or water present, etc.). Data acquisition system 102 may include or be coupled to a thermometer to determine air temperature-related data. Data acquisition system 102 may include or be coupled to a hygrometer to determine humidity-related data. Data acquisition system 102 may include or be coupled to an anemometer to determine wind speed-related data. In this context, the term "coupled" may refer to a direct or indirect connection. The connection may be established, for example, via wired or wireless means.

在各个实施例中,一名或多名听众的群体统计数据可以是或者可以包括与年龄、性别、职业或其任意组合相关的数据。在各个实施例中,数据获取系统102提供的一名或多名听众的群体统计数据可以由数据获取系统102从外部源获得。在各个实施例中,一名或多名听众的群体统计数据可以通过附连于数据获取系统102的数字化形式提供,或者可以作为耦接于(例如,直接耦接或者通过基于云的服务器耦接)数据获取系统102的存储设备上的预加载数据提供。在各个实施例中,一名或多名听众的群体统计数据可以手动提供给数据获取系统102,例如手动键入数据获取系统102,或者根据用户提供的手动指令从远程系统发送至数据获取系统102。在各个实施例中,数据获取系统102可以被设置为自动获得一名或多名听众的群体统计数据,例如,通过有线方式或无线方式从远程系统获得。In various embodiments, the group statistics of one or more listeners may be or may include data related to age, gender, occupation, or any combination thereof. In various embodiments, the group statistics of one or more listeners provided by the data acquisition system 102 may be obtained by the data acquisition system 102 from an external source. In various embodiments, the group statistics of one or more listeners may be provided in digital form attached to the data acquisition system 102, or may be provided as pre-loaded data coupled to (e.g., directly coupled or coupled via a cloud-based server) a storage device coupled to the data acquisition system 102. In various embodiments, the group statistics of one or more listeners may be manually provided to the data acquisition system 102, for example, by manually typing into the data acquisition system 102, or by sending to the data acquisition system 102 from a remote system according to a user-provided manual instruction. In various embodiments, the data acquisition system 102 may be configured to automatically obtain the group statistics of one or more listeners, for example, by wired or wireless means from a remote system.

在各个实施例中,一名或多名听众的心理数据可以是或者可以包括与噪声敏感度(使用例如韦恩斯坦噪声敏感度量表(Weinstein Noise Sensitivity Scale))、感知压力(使用例如科恩感知压力量表(Cohen’s Perceived Stress Scale))、幸福感指数得分(使用例如WHO-5幸福感指数(WHO-5 Well-Being Index))或其任意组合相关的数据。在各个实施例中,数据获取系统102提供的、一名或多名听众的心理数据可以由数据获取系统102从外部源获得。在各个实施例中,一名或多名听众的心理数据可以通过附连于数据获取系统102的数字化形式提供,或者可以作为耦接于(例如,直接耦接或通过基于云的服务器耦接)数据获取系统102的存储设备上的预加载数据提供。在各个实施例中,一名或多名听众的心理数据可以手动提供给数据获取系统102,例如手动键入数据获取系统102,或者根据用户提供的手动指令从远程系统发送至数据获取系统102。在各个实施例中,数据获取系统102可以被设置为自动获得一名或多名听众的心理数据,例如,通过有线方式或无线方式从远程系统获得。In various embodiments, the psychological data of one or more listeners may be, or may include, data related to noise sensitivity (using, for example, the Weinstein Noise Sensitivity Scale), perceived stress (using, for example, the Cohen’s Perceived Stress Scale), well-being index scores (using, for example, the WHO-5 Well-Being Index)), or any combination thereof. In various embodiments, the psychological data of one or more listeners provided by the data acquisition system 102 may be obtained by the data acquisition system 102 from an external source. In various embodiments, the psychological data of one or more listeners may be provided in digital form attached to the data acquisition system 102, or may be provided as pre-loaded data coupled to (e.g., directly coupled or coupled via a cloud-based server) a storage device coupled to the data acquisition system 102. In various embodiments, the psychological data of one or more listeners can be manually provided to the data acquisition system 102, for example, by manually typing it into the data acquisition system 102, or by sending it from a remote system according to a manual instruction provided by the user. In various embodiments, the data acquisition system 102 can be configured to automatically acquire the psychological data of one or more listeners, for example, by acquiring it from a remote system via wired or wireless means.

在各个实施例中,一名或多名听众的生理数据可以是或者可以包括与心率、血压、体温或其任意组合相关的数据。例如,数据获取系统102可以包括或者可以耦接于(例如直接耦接或通过基于云的服务器耦接)一个或多个便携式健康监测设备,该一个或多个便携式健康监测设备被设置为确定或测量一名或多名听众的生理数据。数据获取系统102可以被设置为通过有线方式或无线方式,从一个或多个便携式健康监测设备获得一名或多名听众的生理数据。In various embodiments, the physiological data of one or more listeners may be or may include data related to heart rate, blood pressure, body temperature, or any combination thereof. For example, data acquisition system 102 may include or be coupled to (e.g., directly coupled or coupled via a cloud-based server) one or more portable health monitoring devices configured to determine or measure the physiological data of one or more listeners. Data acquisition system 102 may be configured to acquire the physiological data of one or more listeners from one or more portable health monitoring devices, either wired or wirelessly.

在各个实施例中,感知属性预测器106可以被设置为还基于一个或多个掩蔽增益输入以生成预测。例如,一个或多个掩蔽增益输入可以是一个或多个数字增益级别或一个或多个掩蔽增益波形。In various embodiments, the perceptual attribute predictor 106 may be configured to generate predictions based on one or more masking gain inputs. For example, the one or more masking gain inputs may be one or more digital gain levels or one or more masking gain waveforms.

在各个实施例中,一个或多个预定义感知属性指标可以包括愉悦度、活跃度、平静度、感知响度(perceived loudness)、声音质量(sound quality)、锐度(sharpness)、粗糙度(roughness)或其任意组合。在各个实施例中,一个或多个预定义感知属性指标可以从ISO12913-3中定义的情感质量属性中选择。In various embodiments, one or more predefined perceptual attribute indicators may include pleasantness, activity, calmness, perceived loudness, sound quality, sharpness, roughness, or any combination thereof. In various embodiments, one or more predefined perceptual attribute indicators may be selected from the emotional quality attributes defined in ISO 12913-3.

在各个实施例中,生成的预测可以是或者可以不是单个数值。在各个实施例中,生成的预测可以是非确定性的。在各个其他实施例中,生成的预测可以是确定性的。在各个实施例中,预测可以是概率分布、从概率分布中提取的随机值、一个或多个预定义感知属性的向量或多变量呈现。在各个实施例中,生成的预测可以是预测的属性分布。In various embodiments, the generated prediction may or may not be a single numerical value. In various embodiments, the generated prediction may be nondeterministic. In various other embodiments, the generated prediction may be deterministic. In various embodiments, the prediction may be a probability distribution, random values extracted from a probability distribution, a vector or multivariate representation of one or more predefined perceptual attributes. In various embodiments, the generated prediction may be a predicted attribute distribution.

在各个实施例中,感知属性预测器106可以是或者可以包括概率感知属性预测器。在各个实施例中,感知属性预测器106可以被设置为通过混合、组合或添加周围声景数据和每个掩蔽设置,以生成预测。在各个实施例中,周围声景数据和每个掩蔽设置可以被添加、混合或组合于例如推理引擎(inference engine)中,以在将增强声景数据输入感知属性预测器106之前形成增强声景数据。每个掩蔽设置的掩蔽波形可以与例如掩蔽增益波形的掩蔽增益输入加权(即相乘),以生成加权掩蔽波形。加权掩蔽波形可以与周围声景数据的周围声景波形组合(即添加于周围声景数据的周围声景波形),以生成增强声景波形。感知属性预测器106的预测块可以接收增强声景波形,以生成预测。In various embodiments, the perceptual attribute predictor 106 may be or may include a probabilistic perceptual attribute predictor. In various embodiments, the perceptual attribute predictor 106 may be configured to generate a prediction by mixing, combining, or adding ambient soundscape data and each masking setting. In various embodiments, the ambient soundscape data and each masking setting may be added, mixed, or combined in, for example, an inference engine to form enhanced soundscape data before the enhanced soundscape data is input to the perceptual attribute predictor 106. The masking waveform of each masking setting may be weighted (i.e., multiplied) with, for example, a masking gain waveform, from a masking gain input to generate a weighted masking waveform. The weighted masking waveform may be combined with the ambient soundscape waveform of the ambient soundscape data (i.e., added to the ambient soundscape waveform of the ambient soundscape data) to generate an enhanced soundscape waveform. The prediction block of the perceptual attribute predictor 106 may receive the enhanced soundscape waveform to generate a prediction.

在各个实施例中,感知属性预测器106可以包括设置为提取周围声景数据的特征的声景特征提取器。感知属性预测器106可以包括设置为提取每个掩蔽设置的特征的掩蔽特征提取器。感知属性预测器106可以包括特征级增强器,特征级增强器被设置为基于周围声景数据的提取特征、每个掩蔽设置的提取特征和例如数字增益级别的掩蔽增益输入,生成一个或多个增强声景特征。感知属性预测器106可以包括预测块,预测块被设置为基于一个或多个增强声景特征、周围声景数据的提取特征和每个掩蔽设置的提取特征,生成预测。感知属性预测器106可以被设置为基于周围声景数据的提取特征和每个掩蔽设置的提取特征,生成预测。In various embodiments, the perceptual attribute predictor 106 may include a sound scene feature extractor configured to extract features from the surrounding sound scene data. The perceptual attribute predictor 106 may include a masking feature extractor configured to extract features from each masking setting. The perceptual attribute predictor 106 may include a feature-level enhancer configured to generate one or more enhanced sound scene features based on the extracted features from the surrounding sound scene data, the extracted features from each masking setting, and a masking gain input, such as a digital gain level. The perceptual attribute predictor 106 may include a prediction block configured to generate a prediction based on one or more enhanced sound scene features, the extracted features from the surrounding sound scene data, and the extracted features from each masking setting. The perceptual attribute predictor 106 may be configured to generate a prediction based on the extracted features from the surrounding sound scene data and the extracted features from each masking setting.

在各个实施例中,感知属性预测器106可以包括客观特征块,客观特征块被设置为从周围声景数据和每个掩蔽设置提取客观特征。感知属性预测器106可以被进一步设置为还基于掩蔽增益输入,提取客观特征。感知属性预测器106可以包括主观特征块,主观特征块被设置为从群体统计或情境数据中提取主观特征。主观特征块还可以被设置为处理提取的主观特征和提取的客观特征,以生成预测。感知属性预测器106可以被设置为基于提取的客观特征和提取的主观特征,生成预测。In various embodiments, the perceptual attribute predictor 106 may include an objective feature block configured to extract objective features from surrounding acoustic scene data and each masking setting. The perceptual attribute predictor 106 may be further configured to extract objective features based on masking gain input. The perceptual attribute predictor 106 may include a subjective feature block configured to extract subjective features from crowd statistics or contextual data. The subjective feature block may also be configured to process the extracted subjective features and the extracted objective features to generate a prediction. The perceptual attribute predictor 106 may be configured to generate a prediction based on the extracted objective features and the extracted subjective features.

在各个实施例中,感知属性预测器106可以包括线性回归模型,该线性回归模型被设置为基于声学或心理声学参数,生成预测,该声学或心理声学参数是基于周围声景数据和每个掩蔽设置计算的。In various embodiments, the perceptual attribute predictor 106 may include a linear regression model configured to generate predictions based on acoustic or psychoacoustic parameters calculated based on ambient soundscape data and each masking setting.

在各个实施例中,感知属性预测器106可以包括一个或多个深度神经网络,该一个或多个深度神经网络被设置为基于原始音频、频谱图呈现或原始音频和频谱图呈现的组合,生成预测,该频谱图呈现是基于周围声景数据和每个掩蔽设置计算的。In various embodiments, the perceptual attribute predictor 106 may include one or more deep neural networks configured to generate predictions based on raw audio, spectrogram rendering, or a combination of raw audio and spectrogram rendering, the spectrogram rendering being calculated based on surrounding soundscape data and each masking setting.

在各个实施例中,感知属性预测器106可以包括如本文所述的复合网络或如本文所述的网络和/或模型的任意组合。In various embodiments, the perceptual attribute predictor 106 may include a composite network as described herein or any combination of networks and/or models as described herein.

在各个实施例中,周围声景数据可以由数据获取系统102直接接收、记录、测量或感测,或者可以从外部源接收。在各个实施例中,数据获取系统102可以被设置为基于记录设备、接收设备和/或存储设备的输入,提供或生成周围声景数据。数据获取系统102可以被设置为基于一个或多个收音器(microphones)、一根或多根天线、存储介质或其任意组合的输入,提供或生成周围声景数据。在各个实施例中,数据获取系统102可以被设置为通过直接记录或感测环境的周围声景,以提供或生成周围声景数据。In various embodiments, ambient sound data may be directly received, recorded, measured, or sensed by the data acquisition system 102, or it may be received from an external source. In various embodiments, the data acquisition system 102 may be configured to provide or generate ambient sound data based on input from recording devices, receiving devices, and/or storage devices. The data acquisition system 102 may be configured to provide or generate ambient sound data based on input from one or more microphones, one or more antennas, storage media, or any combination thereof. In various embodiments, the data acquisition system 102 may be configured to provide or generate ambient sound data by directly recording or sensing the ambient sound of the environment.

在各个实施例中,多个掩蔽设置可以包括一个或多个录制的或合成的声音的音轨、一个或多个静音的音轨或者从该一个或多个录制的或合成的声音的音轨和该一个或多个静音的音轨衍生出的一个或多个音轨。In various embodiments, the multiple masking settings may include one or more recorded or synthesized sound tracks, one or more silent sound tracks, or one or more sound tracks derived from the one or more recorded or synthesized sound tracks and the one or more silent sound tracks.

在各个实施例中,播放系统110可以包括一个或多个扬声器、一个或多个虚拟现实或增强现实耳麦、一个或多个耳机、一个或多个耳塞或其任意组合。In various embodiments, the playback system 110 may include one or more speakers, one or more virtual reality or augmented reality headsets, one or more headphones, one or more earbuds, or any combination thereof.

在各个实施例中,可以使用一个或多个计算、处理或电子设备或系统,以实施声景增强系统的各个组件。在一个示例中,数据获取系统102、数据库104、感知属性预测器106、掩蔽设置排序系统108、播放系统110可以在包括收音器和扬声器的单个计算或处理设备或系统中实施,该单个计算或处理设备或系统是例如物理服务器或基于云的服务器、个人计算机系统或移动设备。在另一示例中,数据获取系统102可以通过(具有收音器的)移动设备实施,与此同时,数据库104、感知属性预测器106、掩蔽设置排序系统108和播放系统110可以通过个人计算机系统实施,该个人计算机系统具有直接地或间接地与移动设备无线通信(例如WiFi或蓝牙)的扬声器。在另一示例中,数据获取系统102可以用远程连接于收音器和环境传感器的笔记本电脑实施。数据库104、感知属性预测器106和掩蔽设置排序系统108可以使用通过服务器与笔记本电脑通信的另一计算设备实施。播放系统110可以是与计算机设备无线通信的耳机。In various embodiments, one or more computing, processing, or electronic devices or systems may be used to implement the various components of the soundscape enhancement system. In one example, the data acquisition system 102, database 104, perceptual attribute predictor 106, masking setting sorting system 108, and playback system 110 may be implemented in a single computing or processing device or system including a microphone and speakers, such as a physical server or cloud-based server, personal computer system, or mobile device. In another example, the data acquisition system 102 may be implemented via a mobile device (with a microphone), while the database 104, perceptual attribute predictor 106, masking setting sorting system 108, and playback system 110 may be implemented via a personal computer system having speakers that directly or indirectly communicate wirelessly with the mobile device (e.g., via WiFi or Bluetooth). In yet another example, the data acquisition system 102 may be implemented using a laptop computer remotely connected to a microphone and environmental sensors. The database 104, perceptual attribute predictor 106, and masking setting sorting system 108 may be implemented using another computing device communicating with a laptop computer via a server. The playback system 110 may be headphones that communicate wirelessly with computer equipment.

图2示出根据各个实施例的声景增强系统的形成方法的示意图。该方法可以包括在步骤202提供数据获取系统,该数据获取系统被设置为提供周围声景数据的。该方法还可以包括在步骤204提供包括多个掩蔽设置的数据库。该方法可以进一步包括在步骤206将感知属性预测器耦接于数据获取系统和数据库,该感知属性预测器被设置为基于周围声景数据,为多个掩蔽设置中的每个掩蔽设置,生成在一个或多个预定义感知属性指标上代表感知的预测。该方法可以附加地包括在步骤208提供掩蔽设置排序系统,该掩蔽设置排序系统被设置为基于感知属性预测器生成的预测,确定一个或多个最佳掩蔽设置。该方法可以进一步包括在步骤210提供播放系统,该播放系统被设置为播放或再现一个或多个最佳掩蔽设置。Figure 2 illustrates a schematic diagram of a method for forming a soundscape enhancement system according to various embodiments. The method may include providing a data acquisition system in step 202, the data acquisition system being configured to provide ambient soundscape data. The method may further include providing a database comprising a plurality of masking settings in step 204. The method may further include coupling a perception attribute predictor to the data acquisition system and the database in step 206, the perception attribute predictor being configured to generate a prediction representing perception on one or more predefined perception attribute metrics for each of the plurality of masking settings based on the ambient soundscape data. The method may additionally include providing a masking setting ranking system in step 208, the masking setting ranking system being configured to determine one or more optimal masking settings based on the predictions generated by the perception attribute predictor. The method may further include providing a playback system in step 210, the playback system being configured to play back or reproduce one or more optimal masking settings.

换言之,该方法可以包括将感知属性预测器耦接于数据获取系统和包含多个掩蔽设置的数据库。该方法还可以包括将掩蔽设置排序系统耦接于感知属性预测器,并且将播放系统耦接于掩蔽设置排序系统。In other words, the method may include coupling a perceptual attribute predictor to a data acquisition system and a database containing multiple masking settings. The method may also include coupling a masking setting sorting system to the perceptual attribute predictor and a playback system to the masking setting sorting system.

为避免疑问,图2不旨在限制各个步骤的顺序。例如,步骤202可以在步骤204之前、之后或同时发生。To avoid any doubt, Figure 2 is not intended to restrict the order of the steps. For example, step 202 may occur before, after, or simultaneously with step 204.

在各个实施例中,数据获取系统还可以被设置为提供群体统计或情境数据。感知属性预测器还可以被设置为基于群体统计或情境数据,生成预测。In various embodiments, the data acquisition system may also be configured to provide population statistics or contextual data. The perceived attribute predictor may also be configured to generate predictions based on population statistics or contextual data.

在各个实施例中,群体统计或情境数据可以包括与环境参数相关的数据、一名或多名听众的群体统计数据、一名或多名听众的心理数据、一名或多名听众的生理数据或其任意组合。In various embodiments, group statistics or contextual data may include data related to environmental parameters, group statistics of one or more listeners, psychological data of one or more listeners, physiological data of one or more listeners, or any combination thereof.

在各个实施例中,与环境参数相关的数据可以包括与位置相关的数据、与视觉环境相关的数据、附近的人数、空气温度、湿度、风速或其任意组合。In various embodiments, data related to environmental parameters may include location-related data, visual environment-related data, the number of people nearby, air temperature, humidity, wind speed, or any combination thereof.

在各个实施例中,一名或多名听众的群体统计数据可以包括与年龄、性别、职业或其任意组合相关的数据。In various embodiments, group statistics for one or more listeners may include data related to age, gender, occupation, or any combination thereof.

在各个实施例中,一名或多名听众的心理数据可以包括与噪声敏感度、感知压力、幸福感指数得分或其任意组合相关的数据。In various embodiments, the psychological data of one or more listeners may include data related to noise sensitivity, perceived stress, well-being index scores, or any combination thereof.

在各个实施例中,一名或多名听众的生理数据可以包括与心率、血压、体温或其任意组合相关的数据。In various embodiments, the physiological data of one or more listeners may include data related to heart rate, blood pressure, body temperature, or any combination thereof.

在各个实施例中,感知属性预测器可以被设置为基于一个或多个掩蔽增益输入,生成预测。In various embodiments, the perceptual attribute predictor can be configured to generate predictions based on one or more masking gain inputs.

在各个实施例中,一个或多个预定义感知属性指标可以包括愉悦度、活跃度、事件性、平静度、感知响度、声音质量、锐度、粗糙度或其任意组合。In various embodiments, one or more predefined perceptual attribute metrics may include pleasantness, activity, eventfulness, calmness, perceived loudness, sound quality, sharpness, roughness, or any combination thereof.

在各个实施例中,感知属性预测器可以被设置为通过组合周围声景数据和每个掩蔽设置,以生成预测。In various embodiments, the perception attribute predictor can be configured to generate predictions by combining surrounding soundscape data and each masking setting.

在各个实施例中,感知属性预测器可以包括声景特征提取器,该声景特征提取器被设置为提取周围声景数据的特征。感知属性预测器可以包括掩蔽特征提取器,该掩蔽特征提取器被设置为提取每个掩蔽设置的特征。感知属性预测器可以被设置为基于周围声景数据的提取特征和每个掩蔽设置的提取特征,生成预测。In various embodiments, the perceptual attribute predictor may include a soundscape feature extractor configured to extract features from the surrounding soundscape data. The perceptual attribute predictor may also include a masking feature extractor configured to extract features for each masking setting. The perceptual attribute predictor may be configured to generate a prediction based on the extracted features from the surrounding soundscape data and the extracted features for each masking setting.

在各个实施例中,感知属性预测器可以包括客观特征块,该客观特征块被设置为从周围声景数据和每个掩蔽设置中提取客观特征。感知属性预测器可以包括主观特征块,该主观特征块被设置为从群体统计或情境数据中提取主观特征。感知属性预测器可以被设置为基于提取的客观特征和提取的主观特征,生成预测。In various embodiments, the perceptual attribute predictor may include an objective feature block configured to extract objective features from surrounding acoustic data and each masking setting. The perceptual attribute predictor may also include a subjective feature block configured to extract subjective features from crowd statistics or contextual data. The perceptual attribute predictor may be configured to generate predictions based on the extracted objective and subjective features.

在各个实施例中,感知属性预测器可以包括线性回归模型,该线性回归模型被设置为基于声学或心理声学参数生成预测,该声学或心理声学参数是基于周围声景数据和每个掩蔽设置计算的。In various embodiments, the perceptual attribute predictor may include a linear regression model configured to generate predictions based on acoustic or psychoacoustic parameters calculated based on surrounding acoustic scene data and each masking setting.

在各个实施例中,感知属性预测器可以包括一个或多个深度神经网络,该一个或多个深度神经网络被设置为基于原始音频、频谱图呈现、或原始音频和频谱图呈现的组合生成预测,频谱图呈现是基于周围声景数据和每个掩蔽设置计算的。In various embodiments, the perceptual attribute predictor may include one or more deep neural networks configured to generate predictions based on raw audio, spectrogram rendering, or a combination of raw audio and spectrogram rendering, wherein the spectrogram rendering is calculated based on surrounding soundscape data and each masking setting.

在各个实施例中,数据获取系统可以被设置为基于一个或多个收音器、一根或多根天线、存储介质或其任意组合的输入,提供或生成周围声景数据。In various embodiments, the data acquisition system may be configured to provide or generate ambient sound scene data based on inputs from one or more microphones, one or more antennas, storage media, or any combination thereof.

在各个实施例中,多个掩蔽设置可以包括一个或多个录制的或合成的声音的音轨、一个或多个静音的音轨或从该一个或多个录制的或合成的声音的音轨和该一个或多个静音的音轨衍生出的一个或多个音轨。In various embodiments, the multiple masking settings may include one or more recorded or synthesized sound tracks, one or more muted sound tracks, or one or more sound tracks derived from the one or more recorded or synthesized sound tracks and the one or more muted sound tracks.

在各个实施例中,播放系统可以包括一个或多个扬声器、一个或多个虚拟现实或增强现实耳麦、一个或多个耳机、一个或多个耳塞或其任意组合。In various embodiments, the playback system may include one or more speakers, one or more virtual reality or augmented reality headsets, one or more headphones, one or more earbuds, or any combination thereof.

在各个实施例中,生成的预测可以是非确定性的(non-deterministic)。In various embodiments, the generated predictions may be non-deterministic.

图3示出根据各个实施例的声景增强系统的示意图。声景增强系统也可以称为自动掩蔽选择系统(automatic masker selection system,AMSS)。声景增强系统可以包括数据获取系统302,数据获取系统302负责提供周围声景数据312,即声学环境的音频和/或直接从声学环境的音频中计算的参数。为从声学环境获得周围声景数据312,数据获取系统302可以使用例如一个或多个收音器的实时记录、从无线天线接收的数字音频信号、存储于附连的存储介质上的数据或其任意组合。Figure 3 illustrates a schematic diagram of a soundscape enhancement system according to various embodiments. The soundscape enhancement system may also be referred to as an automatic masker selection system (AMSS). The soundscape enhancement system may include a data acquisition system 302, which is responsible for providing ambient soundscape data 312, i.e., audio of the acoustic environment and/or parameters calculated directly from the audio of the acoustic environment. To obtain the ambient soundscape data 312 from the acoustic environment, the data acquisition system 302 may use, for example, real-time recordings from one or more microphones, digital audio signals received from a wireless antenna, data stored on an attached storage medium, or any combination thereof.

此外,声景增强系统可以包括数据库,该数据库包括多个掩蔽设置304(替代地称为候选掩蔽设置组)。术语“掩蔽设置”可以指可通过任何形式的播放添加于312中的声学环境的、与音轨的预期播放所需的其他参数或效果相结合的任何音轨。掩蔽设置可以替代地称为“掩蔽器”。该参数或效果可以包括例如数字增益级别、空间化(spatialization)、动态包络(dynamic envelope)和/或滤波(filtering)的应用。掩蔽设置的示例可以包括但不限于由以不同音量播放的录制的或合成的声音组成的音轨、仅由静音组成的音轨和/或从304中的一个或多个其他音轨的函数中衍生的音轨,该仅由静音组成的音轨的播放不会对312中的声学环境构成任何声音添加。掩蔽设置可以替代地或附加地指选择掩蔽器的过程,例如从音轨的数据库304中选择掩蔽器的过程。掩蔽设置策略可以是自主的(autonomous),并且可以(通过例如监控收音器或通过数字化输入)适应当前声景的动态(例如振幅、频率内容、情境)。Furthermore, the soundscape enhancement system may include a database comprising multiple masking settings 304 (alternatively referred to as candidate masking setting groups). The term "masking setting" can refer to any audio track that can be added to the acoustic environment in 312 through any form of playback, in combination with other parameters or effects required for the intended playback of the audio track. Masking settings can also be referred to as "masks." These parameters or effects may include, for example, the application of digital gain levels, spatialization, dynamic envelope, and/or filtering. Examples of masking settings may include, but are not limited to, audio tracks consisting of recorded or synthesized sounds played at different volumes, audio tracks consisting only of silence, and/or audio tracks derived from functions of one or more other audio tracks in 304, the playback of which constitutes no sound addition to the acoustic environment in 312. Masking settings can alternatively or additionally refer to the process of selecting a mask, such as the process of selecting a mask from the database of audio tracks 304. Masking strategies can be autonomous and can adapt to the dynamics of the current soundscape (e.g., amplitude, frequency content, context) by means of monitoring the microphone or by digital input.

数据获取系统302还可以可选地负责采集群体统计或情境数据314作为声学环境本身的辅助参数,并且负责提供群体统计或情境数据314。该群体统计或情境数据314可以包括但不限于(1)非声学环境参数,例如位置(通过GPS传感器采集)、附近人数(通过摄像头采集)、空气温度(通过温度计采集)、湿度(通过湿度计采集)和/或风速(通过风速计采集);(2)声学环境中真实听众或假想听众的群体统计数据(例如,通过附连于系统的数字化形式(digital form)采集或者作为存储设备上的预加载数据),例如年龄、性别和/或职业;(3)声学环境中真实听众或假想听众的心理数据(通过例如与群体统计方法相似的方法采集),例如噪声敏感度(使用例如韦恩斯坦噪声敏感度量表获得)、感知压力(使用例如科恩感知压力量表获得),以及/或者幸福感指数得分(使用例如WHO-5幸福感指数获得);和/或声学环境中真实听众或假想听众的生理数据(通过例如便携式健康监测设备获得),例如心率、血压和/或体温。The data acquisition system 302 may also optionally be responsible for collecting group statistics or contextual data 314 as auxiliary parameters of the acoustic environment itself, and for providing group statistics or contextual data 314. The group statistics or contextual data 314 may include, but are not limited to, (1) non-acoustic environmental parameters, such as location (collected by a GPS sensor), number of people nearby (collected by a camera), air temperature (collected by a thermometer), humidity (collected by a hygrometer) and/or wind speed (collected by an anemometer); (2) group statistics of real or hypothetical listeners in the acoustic environment (e.g., collected by a digital form attached to the system or as pre-loaded data on a storage device), such as age, gender and/or occupation; (3) psychological data of real or hypothetical listeners in the acoustic environment (collected by, for example, methods similar to group statistics), such as noise sensitivity (obtained using, for example, the Weinstein Noise Sensitivity Scale), perceived stress (obtained using, for example, the Cohen Perceived Stress Scale), and/or well-being index scores (obtained using, for example, the WHO-5 Well-being Index); and/or physiological data of real or hypothetical listeners in the acoustic environment (obtained by, for example, a portable health monitoring device), such as heart rate, blood pressure and/or body temperature.

数据获取系统302可以将周围声景数据312和可选的群体统计或情境数据314提供给感知属性预测器306。在给定周围声景数据312、数据库304的掩蔽设置以及可选的群体统计或情境数据314作为输入的情况下,感知属性预测器306可以为数据库304中的每个掩蔽设置,输出在一个或多个预定义感知属性指标上的预测(对314中的数据所对应的听众的预测),如同每个掩蔽设置的播放将在声学环境中实现。例如,一个或多个预定义感知属性指标可以源自用于数据库304中的每个掩蔽设置的ISO12913-3的圆周八分指标。例如,一个或多个预定义感知属性指标可以基于用于数据库304中每个掩蔽设置的愉悦度、事件性、声音质量、感知响度、锐度、粗糙度等。图4示出ISO12913-3中呈现的情感质量属性的二维圆周八分模型,根据各个实施例的感知属性预测器306的预测可以基于该二维圆周八分模型。The data acquisition system 302 can provide ambient soundscape data 312 and optional crowd statistics or contextual data 314 to the perceptual attribute predictor 306. Given the ambient soundscape data 312, the masking settings of the database 304, and the optional crowd statistics or contextual data 314 as input, the perceptual attribute predictor 306 can output a prediction on one or more predefined perceptual attribute metrics for each masking setting in the database 304 (a prediction of the audience corresponding to the data in 314), as if the playback of each masking setting were to occur in the acoustic environment. For example, one or more predefined perceptual attribute metrics can be derived from the ISO 12913-3 octet metrics used for each masking setting in the database 304. For example, one or more predefined perceptual attribute metrics can be based on pleasantness, eventfulness, sound quality, perceived loudness, sharpness, roughness, etc., used for each masking setting in the database 304. Figure 4 illustrates a two-dimensional octet model of emotional quality attributes presented in ISO 12913-3, and the predictions of the perceptual attribute predictor 306 according to various embodiments can be based on this two-dimensional octet model.

在图3中,该预测的集合由316表示。预测316可以是或者可以不是单个数值。类似地,在各个实施例中,生成的预测316可以是非确定性的,与此同时,在各个其他实施例中,生成的预测316可以是确定性的。术语“确定性的”可以指,每当周围声景数据312、数据库304的多个掩蔽设置以及可选的群体统计或情境数据314的相同输入被提供给感知属性预测器306时,就会产生相同的输出。概率分布(或从概率分布中抽取的随机值)、向量或感知属性的任何其他多变量呈现可以是作为预测316的可能输出。In Figure 3, the set of predictions is represented by 316. Prediction 316 may or may not be a single numerical value. Similarly, in various embodiments, the generated prediction 316 may be nondeterministic, while in various other embodiments, the generated prediction 316 may be deterministic. The term "deterministic" can mean that the same output is produced whenever the same inputs of ambient soundscape data 312, multiple masking settings of database 304, and optional crowd statistics or contextual data 314 are provided to the perceptual attribute predictor 306. Probability distributions (or random values drawn from probability distributions), vectors, or any other multivariate representation of perceptual attributes can be possible outputs of prediction 316.

感知属性预测器306可以是或者可以包括从周围声景数据312、数据库304的多个掩蔽设置以及可选的群体统计或情境数据314中获取相同的输入以输出预测316的任何模型。该模型可以在将预测316作为输出输送之前,内部转换周围声景数据312、数据库304的多个掩蔽设置以及群体统计或情境数据314。该模型的示例可以包括但不限于:(1)线性回归模型,该线性回归模型获取从周围声景数据312计算的声学或心理声学参数、数据库304的多个掩蔽设置以及从群体统计或情境数据314计算的聚合平均值,以输出预测316;(2)深度神经网络,该深度神经网络获取从周围声景数据312中计算的呈现、数据库304的多个掩蔽设置,其中串联有群体统计或情境数据314的呈现,以输出预测316;(3)如图5至图7所示的复合网络;和/或(4)如本文所述的任意组合。The perceptual attribute predictor 306 may be or may include any model that takes the same input from ambient soundscape data 312, multiple masking settings of database 304, and optional crowd statistics or contextual data 314 to output prediction 316. The model may internally transform the ambient soundscape data 312, the multiple masking settings of database 304, and the crowd statistics or contextual data 314 before delivering prediction 316 as output. Examples of the model may include, but are not limited to: (1) a linear regression model that takes acoustic or psychoacoustic parameters calculated from ambient soundscape data 312, multiple masking settings of database 304, and aggregated averages calculated from crowd statistics or contextual data 314 to output prediction 316; (2) a deep neural network that takes presentations calculated from ambient soundscape data 312, multiple masking settings of database 304, wherein presentations of crowd statistics or contextual data 314 are concatenated, to output prediction 316; (3) a composite network as shown in Figures 5 through 7; and/or (4) any combination as described herein.

图5是根据各个实施例的感知属性预测器506的示意图。在各个实施例中,感知属性预测器506可以被设置为通过将周围声景数据和每个掩蔽设置组合,以生成预测,例如,在推理引擎中以预校准的声景与掩蔽器的比率(soundscape-to-masker ratios)进行组合。每个掩蔽设置的掩蔽波形504a可以与例如掩蔽增益波形的掩蔽增益输入504b加权(即相乘),以生成加权掩蔽波形。加权掩蔽波形可以与周围声景数据的周围声景波形512组合(即添加于周围声景数据的周围声景波形512),以生成增强声景波形518。感知属性预测器506的预测块520可以接收增强声景波形518,以生成预测510。Figure 5 is a schematic diagram of a perceptual attribute predictor 506 according to various embodiments. In various embodiments, the perceptual attribute predictor 506 may be configured to generate predictions by combining ambient soundscape data and each masking setting, for example, by combining them in an inference engine with pre-calibrated soundscape-to-masker ratios. The masking waveform 504a of each masking setting may be weighted (i.e. multiplied) with a masking gain input 504b, such as a masking gain waveform, to generate a weighted masking waveform. The weighted masking waveform may be combined with the ambient soundscape waveform 512 of the ambient soundscape data (i.e., added to the ambient soundscape waveform 512 of the ambient soundscape data) to generate an enhanced soundscape waveform 518. The prediction block 520 of the perceptual attribute predictor 506 may receive the enhanced soundscape waveform 518 to generate a prediction 510.

在将增强声景波形输入感知属性预测器中之前,各个其他实施例可以不需要混合、组合或添加周围声景数据和每个加权掩蔽波形以形成增强声景波形。在将增强声景波形输入感知属性预测器之前不需要通过混合、组合或添加周围声景波形和每个加权掩蔽波形以形成增强声景数据,各个实施例可以在实际部署中允许实现较高效的计算和带宽系统。Various other embodiments may not require mixing, combining, or adding ambient sound scene data and each weighted masking waveform to form the enhanced sound scene waveform before inputting the enhanced sound scene waveform into the perceptual attribute predictor. By eliminating the need to mix, combine, or add ambient sound scene waveforms and each weighted masking waveform to form the enhanced sound scene data before inputting the enhanced sound scene waveform into the perceptual attribute predictor, these embodiments may allow for more efficient computational and bandwidth systems in practical deployments.

图6是根据各个实施例的感知属性预测器606的示意图。在各个实施例中,感知属性预测器606可以包括声景特征提取器622,声景特征提取器622被设置为提取周围声景数据的特征624,即提取周围声景数据的周围声景波形612的特征624。感知属性预测器606可以包括掩蔽特征提取器626,掩蔽特征提取器626被设置为提取每个掩蔽设置的特征628,即提取每个掩蔽设置的掩蔽波形604a的特征628。感知属性预测器606可以包括特征级别增强器630,特征级别增强器630被设置为基于周围声景数据的提取特征624、每个掩蔽设置的提取特征628和例如数字增益级别的掩蔽增益输入604b,生成一个或多个增强声景特征632。感知属性预测器606可以包括预测块620,预测块620被设置为基于一个或多个增强声景特征632、周围声景数据的提取特征624和每个掩蔽设置的提取特征628,生成预测610。概括而言,感知属性预测器606可以被设置为基于周围声景数据的提取特征624和每个掩蔽设置的提取特征628,生成预测610。Figure 6 is a schematic diagram of a perceptual attribute predictor 606 according to various embodiments. In various embodiments, the perceptual attribute predictor 606 may include a soundscape feature extractor 622, which is configured to extract features 624 of the surrounding soundscape data, specifically features 624 of the surrounding soundscape waveform 612. The perceptual attribute predictor 606 may include a masking feature extractor 626, which is configured to extract features 628 for each masking setting, specifically features 628 of the masking waveform 604a for each masking setting. The perceptual attribute predictor 606 may include a feature level enhancer 630, which is configured to generate one or more enhanced soundscape features 632 based on the extracted features 624 of the surrounding soundscape data, the extracted features 628 of each masking setting, and a masking gain input 604b, such as a digital gain level. The perceptual attribute predictor 606 may include a prediction block 620, which is configured to generate a prediction 610 based on one or more enhanced soundscape features 632, extracted features 624 of the surrounding soundscape data, and extracted features 628 of each masking setting. In summary, the perceptual attribute predictor 606 may be configured to generate a prediction 610 based on extracted features 624 of the surrounding soundscape data and extracted features 628 of each masking setting.

图7是根据各个实施例的感知属性预测器706的示意图。在各个实施例中,感知属性预测器706可以包括客观特征块734,客观特征块734被设置为从周围声景数据和每个掩蔽设置中提取客观特征736,即从周围声景数据的周围声景波形712和每个掩蔽设置的掩蔽波形704a中提取客观特征736。感知属性预测器706可以被进一步设置为还基于掩蔽增益输入704b,提取客观特征736。感知属性预测器706可以包括主观特征块738,主观特征块738被设置为从群体统计或情境数据714中提取主观特征。主观特征块738还可以被设置为处理提取的主观特征和提取的客观特征736,以生成预测710。感知属性预测器706可以被设置为基于提取的客观特征736和提取的主观特征,生成预测710。Figure 7 is a schematic diagram of a perceptual attribute predictor 706 according to various embodiments. In various embodiments, the perceptual attribute predictor 706 may include an objective feature block 734, which is configured to extract objective features 736 from the surrounding soundscape data and each masking setting, specifically from the surrounding soundscape waveform 712 of the surrounding soundscape data and the masking waveform 704a of each masking setting. The perceptual attribute predictor 706 may be further configured to extract the objective features 736 based on the masking gain input 704b. The perceptual attribute predictor 706 may include a subjective feature block 738, which is configured to extract subjective features from crowd statistics or contextual data 714. The subjective feature block 738 may also be configured to process the extracted subjective features and the extracted objective features 736 to generate a prediction 710. The perceptual attribute predictor 706 may be configured to generate a prediction 710 based on the extracted objective features 736 and the extracted subjective features.

再参考图3,感知属性预测器306提供的预测316可以输入掩蔽设置排序系统308中,以生成单个最佳掩蔽设置340。掩蔽设置排序系统308可以通过任何合适的指标确定最佳性(optimality),该指标可以是但不限于感知响度、声音质量、锐度、粗糙度、图4所示的一个或多个情感质量属性和/或其任意组合中的最大值或最小值。替代地,可以设想感知属性预测器306提供的预测316可以输入掩蔽设置排序系统308中,以生成多个最佳掩蔽设置。对于给定的指标,可以有多个最佳掩蔽设置(例如,如果用户决定的指标中绑定有多个掩蔽设置)。Referring again to Figure 3, the prediction 316 provided by the perceptual attribute predictor 306 can be input into the masking setting ranking system 308 to generate a single optimal masking setting 340. The masking setting ranking system 308 can determine the optimality using any suitable metric, which can be, but is not limited to, the maximum or minimum value of perceptual loudness, sound quality, sharpness, roughness, one or more emotional quality attributes shown in Figure 4, and/or any combination thereof. Alternatively, it is conceivable that the prediction 316 provided by the perceptual attribute predictor 306 can be input into the masking setting ranking system 308 to generate multiple optimal masking settings. For a given metric, there can be multiple optimal masking settings (e.g., if multiple masking settings are bound to the metric determined by the user).

最佳掩蔽设置340可以用作播放系统310的输入。播放系统310可以实现最佳掩蔽设置340在声学环境中的实际播放或再现。播放系统310可以包括但不限于一个或多个扬声器、一个或多个虚拟现实或增强现实耳麦、一个或多个耳机、一个或多个耳塞和/或其任意组合。The optimal masking setting 340 can be used as input to the playback system 310. The playback system 310 can realize the actual playback or reproduction of the optimal masking setting 340 in an acoustic environment. The playback system 310 may include, but is not limited to, one or more speakers, one or more virtual reality or augmented reality headsets, one or more headphones, one or more earbuds, and/or any combination thereof.

各个实施例可以用于具有较差的声音质量的室外城市区域(例如高速公路旁的公园),以减轻周围声学环境中所感知的、噪声的有害影响,或者用于室内区域(例如疗养院内部),以改善周围声学环境的声学舒适度和整体质量。由于该系统有效地旨在改变周围声学环境的感知,拥有室内或室外区域的任何商业实体可能对该系统感兴趣,该室内或室外区域的声学环境需要保持于固定的或期望的条件下。Various embodiments can be used in outdoor urban areas with poor sound quality (e.g., parks next to highways) to mitigate the harmful effects of perceived noise in the surrounding acoustic environment, or in indoor areas (e.g., inside nursing homes) to improve acoustic comfort and overall quality of the surrounding acoustic environment. Because the system is effectively designed to alter the perception of the surrounding acoustic environment, any commercial entity with indoor or outdoor areas whose acoustic environment needs to be maintained under fixed or desired conditions may be interested in the system.

各个实施例能够选择掩蔽器且将掩蔽器添加于现实的周围声景和假想的周围声景中。如果将任意音频输入数据获取系统以代替收音器的实时录音,则各个实施例还可以通过添加掩蔽器,用以创建抽象的“体验区”。任何有兴趣基于听众感知制作或修改虚拟现实、混合现实或增强现实中的声景的商业实体都会对这一点感兴趣。Various embodiments are capable of selecting and adding masks to both real and imaginary ambient soundscapes. If an arbitrary audio input data acquisition system is used in place of a microphone for real-time recording, the embodiments can also create abstract “experience zones” by adding masks. Any commercial entity interested in creating or modifying soundscapes in virtual reality, mixed reality, or augmented reality based on audience perception will be of interest in this.

各个实施例可以涉及系统,该系统基于优化预测的感知响应(例如,感知响度、ISO12913-3的圆周八分指标——愉悦度、烦扰度、安宁度、平静度、活跃度、事件性等),将待引入现有声景的音轨(即掩蔽器)(通过例如扬声器播放或通过例如耳机进行数字化添加)选择且设置于混合/增强声景。Various embodiments may involve systems that, based on optimized predicted perceptual responses (e.g., perceived loudness, ISO 12913-3 octet indexes—pleasantness, annoyance, tranquility, calmness, activity, eventfulness, etc.), select and set audio tracks (i.e., maskers) to be introduced into an existing soundscape (either played through, for example, speakers or digitally added through, for example, headphones) in a mixed/enhanced soundscape.

各个实施例可以基于感知属性预测器对人类感知的预测,选择掩蔽设置或掩蔽器以添加于周围声学环境,该感知属性预测器根据周围声景数据(现实的或假想的周围声景数据)、候选掩蔽设置的数据以及可选的该周围声景的听众的群体统计或情境数据,输出感知属性指标的值或分布。这体现了“声景”的ISO12913的定义,即将声学环境视为由一个人或多个人在情境中所感知的声学环境。相比之下,现有方法可能基于周围声学环境的客观声学参数(例如频谱、声压级、声源类型和方向),选择掩蔽设置或掩蔽器,并且可能不考虑人类感知对掩蔽选择或掩蔽设置选择的主观影响。Various embodiments can select masking settings or maskers to be added to the surrounding acoustic environment based on predictions of human perception by a perceptual attribute predictor. This perceptual attribute predictor outputs values or distributions of perceptual attribute indicators based on surrounding soundscape data (real or hypothetical), data on candidate masking settings, and group statistics or contextual data of the audience for the possible surrounding soundscape. This embodies the ISO 12913 definition of a "soundscape," which considers the acoustic environment as perceived by one or more people within a context. In contrast, existing methods may select masking settings or maskers based on objective acoustic parameters of the surrounding acoustic environment (e.g., spectrum, sound pressure level, sound source type, and direction), and may not consider the subjective influence of human perception on masking selection or masking setting selection.

在各个实施例中,感知属性预测器可以近乎实时地输出感知预测(输入与预测之间的延迟小于10秒),并且因此可以自动实时地为随时间改变的声学环境,提出随时间改变的最佳掩蔽设置的建议。换言之,当声学环境随时间改变时,各个实施例可以自主地改变正在播放的掩蔽器(即音轨)和/或设置(例如增益级别、空间化)。相比之下,现有系统仅允许静态掩蔽设置(仅潜在地应用有动态音量/滤波器效果)或用户控制的掩蔽设置(即,一个人必须基于其自己的个人/专家对最优性的意见,手动改变设置)。In various embodiments, the perceptual attribute predictor can output perceptual predictions in near real-time (with a delay of less than 10 seconds between input and prediction), and therefore can automatically and in real-time suggest optimal masking settings for a time-varying acoustic environment. In other words, as the acoustic environment changes over time, the embodiments can autonomously change the masking (i.e., audio track) and/or settings (e.g., gain level, spatialization) being played. In contrast, existing systems only allow static masking settings (potentially applying only dynamic volume/filter effects) or user-controlled masking settings (i.e., one must manually change the settings based on their own personal/expert opinion on optimality).

在各个实施例中,用户可以将系统设置为在多个选择方案中自动选取最佳掩蔽器。选择方案可以是任意的,但是各个实施例可以允许常规解决方案所没有的以下选择方案:(1)概率top-k:对不同掩蔽器中的组合掩蔽器和现有声景的预测感知响应值进行排序,并且在给出前(或后)k个排序值的情况下,从k个掩蔽设置中随机选取一个掩蔽设置;(2)概率分布:从系统生成的概率分布中抽取组合掩蔽器和现有声景的感知响应值的预测作为随机样本。In various embodiments, the user can configure the system to automatically select the best mask from multiple options. The selection options can be arbitrary, but various embodiments may allow the following selection options not available in conventional solutions: (1) probabilistic top-k: sorting the predicted perceptual response values of the combined mask and the existing soundscape in different masks, and randomly selecting a mask setting from the k mask settings given the top (or bottom) k sorting values; (2) probabilistic distribution: extracting the predicted perceptual response values of the combined mask and the existing soundscape as random samples from the probability distribution generated by the system.

现有系统和掩蔽选择策略可能不允许通过动态策略选择最佳掩蔽器,并且可能具有以下特征中的一个特征:(1)手动选择:系统用户或管理员基于其自己的意见选择最佳掩蔽器;(2)确定性:从不同掩蔽器中选取最佳掩蔽器作为组合的掩蔽器和现有声景的感知响应值的确定性的、预定义的函数。Existing systems and masking selection strategies may not allow for the selection of the best mask through dynamic strategies and may have one of the following characteristics: (1) manual selection: system users or administrators select the best mask based on their own opinions; (2) determinism: the selection of the best mask from different masks as a deterministic, predefined function of the combined mask and the perceived response value of the existing soundscape.

各个实施例可以包括一个或多个音频换能器(例如扬声器、耳机)和一个或多个传感器(例如收音器、环境传感器),该一个或多个音频换能器和该一个或多个传感器相互耦接于一个或多个计算设备。各个实施例可以具有现有系统/解决方案所不具备的优势。Various embodiments may include one or more audio transducers (e.g., speakers, headphones) and one or more sensors (e.g., microphones, environmental sensors), wherein the audio transducers and sensors are coupled to one or more computing devices. Various embodiments may offer advantages not available in existing systems/solutions.

例如,各个实施例可以采用非能量掩蔽。能量掩蔽利用人类听觉外周过程(peripheral human auditory processes)的物理限制。掩蔽声音的振幅和频率带宽被调整以使得目标噪声在内耳处是完全地或部分地无法被听到的(inaudible)。因此,通常采用例如声压级或心理声学参数的客观指标,以优化或选择能量掩蔽器。各个实施例可以参考(但不限于)ISO12913-1中定义的声景概念。在文献中,使用附加声音进行的声景增强通常也被称为掩蔽。掩蔽技术可以包括能量掩蔽和信息掩蔽(或感知掩蔽),其中信息掩蔽由与声音的可察觉性的影响构成,该声音的可察觉性与较高级的大脑中心中的音频(audio inthe higher brain centres)相关。因此,掩蔽器可以由感知因素(例如ISO12913-2中的感知情感属性、感知的烦扰度)决定,并且不是仅基于现有系统/解决方案中的客观指标。For example, various embodiments may employ non-energy masking. Energy masking utilizes the physical limitations of peripheral human auditory processes. The amplitude and frequency bandwidth of the masking sound are adjusted so that the target noise is completely or partially inaudible at the inner ear. Therefore, objective metrics such as sound pressure level or psychoacoustic parameters are typically used to optimize or select the energy masker. Various embodiments may refer to (but are not limited to) the soundscape concept as defined in ISO 12913-1. In the literature, soundscape enhancement using additional sound is also commonly referred to as masking. Masking techniques may include energy masking and information masking (or perceptual masking), where information masking consists of effects related to the perceptibility of sound, which is associated with audio in the higher brain centers. Therefore, the mask can be determined by perceived factors (such as perceived emotional attributes and perceived annoyance in ISO 12913-2) and is not based solely on objective indicators in existing systems/solutions.

非能量掩蔽的优势可以在关于使用自然声音进行的、交通噪声的掩蔽的文献中得到证实,在该文献中,随着目标噪声级别的增加,能量掩蔽的感知响度的降低逐渐消失(nullified)。这与在较高级别的交通噪声下,提供信息掩蔽的掩蔽器以低最多6dB的级别呈现时,统计上显著的感知响度的降低和声景质量改善相反。此外,在使用鸟鸣声作为掩蔽器的情况下,也观察到相似的感知响度的显著降低和声景质量的改善,其中使用鸟鸣声作为掩蔽器无法在能量上掩蔽交通噪声,这进一步证明非能量掩蔽的有效性。The advantages of non-energy masking are evidenced in literature on traffic noise masking using natural sounds, where the reduction in perceived loudness from energy masking gradually disappears as the target noise level increases. This contrasts with the statistically significant reduction in perceived loudness and improvement in soundscape quality observed when the masking device providing information masking is presented at a level up to 6 dB lower at higher levels of traffic noise. Furthermore, similar significant reductions in perceived loudness and improvements in soundscape quality were observed when using birdsong as a masking device, where birdsong as a masking device fails to energy mask traffic noise, further demonstrating the effectiveness of non-energy masking.

通过辅助输入将音频输入情境化的能力可以允许各个实施例为任意目标声景(以代替仅限于现有文献中描述的办公环境)选择掩蔽器。根据各个实施例的掩蔽器也可以不限于从随机噪声生成的音轨,也不局限于现有文献中描述的自然环境的声音(naturalsounds)。只要所期望出现的感知属性被优化(优化可以指“最大化”或“最小化”),可以添加的掩蔽器的数量就可以不受限制。The ability to contextualize audio input via auxiliary input allows various embodiments to select a mask for any target soundscape (instead of being limited to office environments described in existing literature). Masks according to various embodiments may also be not limited to audio tracks generated from random noise, nor to natural sounds as described in existing literature. The number of masks that can be added is unlimited, as long as the desired perceptual properties are optimized (optimization can mean "maximizing" or "minimizing").

在各个实施例中,音频输入数据可以是不限格式的,并且可以是物理收音器的实时数据或从存储介质访问的音轨。类似地,根据各个实施例的掩蔽器可以流式传输(streamed)至扬声器(或耳机)或与存储介质的目标声景进行数字化混合,并且在虚拟环境中作为单个组合音轨播放。因此,各个实施例能够掩蔽现实世界环境以及虚拟环境中的声景(例如在元宇宙应用程序中)。In various embodiments, the audio input data can be of any format and can be real-time data from a physical microphone or an audio track accessed from a storage medium. Similarly, the masking according to various embodiments can be streamed to speakers (or headphones) or digitally mixed with a target soundscape on a storage medium and played as a single combined audio track in a virtual environment. Therefore, the various embodiments are capable of masking soundscapes in both real-world and virtual environments (e.g., in metaverse applications).

现有文献中描述的动态掩蔽是基于客观指标的,例如声级、音频频谱,或者是通过与所识别的目标噪声源的相近度实现的。各个实施例可以基于预测的感知属性动态地或适应地调整掩蔽设置。掩蔽设置可以不限于声级或频谱调整。Dynamic masking described in existing literature is based on objective metrics, such as sound level, audio spectrum, or proximity to the identified target noise source. Various embodiments can dynamically or adaptively adjust the masking settings based on predicted perceptual properties. The masking settings are not limited to sound level or spectrum adjustments.

示例研究1Example Study 1

介绍introduce

已知仅基于给定声学环境的声压级(sound pressure level,SPL)的指示通常不足以反映过度噪声造成的烦扰度级别和对生活质量的影响。这已经引起了噪声控制的声景方法的兴起,该方法着重于改善噪声感知属性的干预措施,以代替简单地降低SPL。关于声景的国际标准ISO12913系列旨在通过提供愉悦度和事件性的圆周模型以规范该方法,可以基于对周围声学环境的主观评估,根据该圆周模型比较干预措施。It is known that indications based solely on the sound pressure level (SPL) of a given acoustic environment are often insufficient to reflect the level of annoyance caused by excessive noise and its impact on quality of life. This has led to the rise of the soundscape approach to noise control, which focuses on interventions that improve perceived noise properties rather than simply reducing SPL. The ISO 12913 series of international standards on soundscapes aims to standardize this approach by providing a circular model of pleasantness and eventability, which allows for the comparison of interventions based on subjective assessments of the surrounding acoustic environment.

因此,许多研究已经利用声景增强技术,通过对城市或室内声学环境添加掩蔽器,以改变声景的感知,从而优化例如主观评价的愉悦度或平静度的指标。然而,掩蔽器的选择通常是任意的、专家指导的或是基于事后分析的。掩蔽器的任意选择可能在实现所期望出现的感知改变的方面是不可靠的,专家指导的选择是耗力且耗时的,并且事后分析可能无法泛化至在没有观察到的情境中的、没有看到的掩蔽器和声景。Therefore, many studies have utilized soundscape enhancement techniques to modify the perception of soundscapes by adding maskers to urban or indoor acoustic environments, thereby optimizing metrics such as subjective evaluations of pleasantness or calmness. However, the selection of maskers is often arbitrary, expert-guided, or based on post-hoc analysis. Arbitrary masker selection may be unreliable in achieving the desired perceptual changes, expert-guided selection is laborious and time-consuming, and post-hoc analysis may not generalize to unseen maskers and soundscapes in unobserved contexts.

克服这些限制的方法是根据声景和掩蔽器的、声学上的多样化选择,训练预测模型以预测如被给予原始听觉刺激的人主观评估的、某个或某种感知属性的值。训练完成后,掩蔽器对没有看到的声景的模拟添加可以作为对模型的输入馈送,以获得所述感知属性的预测。然后,可以选择对属性产生最多提升或最多降低的掩蔽器作为最佳掩蔽器,以用于实时增强。One approach to overcome these limitations is to train a predictive model to predict the value of one or more perceptual attributes, as subjectively assessed by a person given the original auditory stimulus, based on a diverse selection of acoustically varied soundscapes and maskers. After training, simulated additions of unseen soundscapes to the masker can be fed into the model as input to obtain predictions of the perceptual attributes. The masker that produces the most enhancement or reduction to the attribute can then be selected as the optimal mask for real-time augmentation.

各个实施例可以涉及用于自动掩蔽选择系统(automatic masker selectionsystem,AMSS)的、通过优化声学环境的愉悦度以增强城市声景的神经学方法(neuralapproach)。通过使用概率输出方案,模型可以为每个声景与掩蔽器的组合预测愉悦度分布,以代替单个确定值,这允许明确获取预测的置信度以及预测的愉悦度。Various embodiments may relate to a neural approach for an automatic masker selection system (AMSS) that enhances urban soundscapes by optimizing the pleasantness of the acoustic environment. By using a probabilistic output scheme, the model can predict the pleasantness distribution for each combination of soundscape and masker, instead of a single deterministic value. This allows for explicit acquisition of the confidence level of the prediction as well as the predicted pleasantness.

相关工作Related work

在声景研究中,开发用于感知属性的预测模型的研究主要集中于较简单的机器学习模型上,例如线性回归、支持向量机(support vector machines,SVM)和浅层多级感知器(shallow multilayer perceptrons)。机器学习模型使用基于声学测量、心理声学参数、环境特征或由主成分分析阐明的机器学习模型的某个或某种线性组合的输入特征。还提出了用于李克特量表(Likert scales)的标度指标(scaling metric),以解释用于开发这些系统的量表中的非线性。In soundscape research, studies developing predictive models for perceptual attributes have primarily focused on simpler machine learning models, such as linear regression, support vector machines (SVMs), and shallow multilayer perceptrons. These machine learning models utilize input features based on acoustic measurements, psychoacoustic parameters, environmental features, or one or more linear combinations of machine learning models elucidated by principal component analysis. Scaling metrics for Likert scales have also been proposed to account for nonlinearities in the scales used to develop these systems.

另一方面,最近的系统性综述示出,尽管深度神经网络在例如声音事件定位、检测和分类的较“客观”的任务中很普遍,但是利用深度神经网络预测声景感知属性值的研究却很少。就此而言,最重要的研究似乎是比较SVM、卷积神经网络(convolutional neuralnetwork,CNN)、长短期记忆网络(long short-term memory network,LSTM)和微调VGGish网络(fine-tuned VGGish network)在情绪-声景数据集(Emo-Soundscapes dataset)上的表现的研究。根据通过自我评估人偶模型(Self-Assessment Manikin)进行的主观配对比较,情绪-声景包含1213个按效价(愉悦度)和唤醒度(事件性)排序的音频片段。微调的VGGish模型和从零训练的CNN在预测音频片段的效价和唤醒度方面,分别呈现为具有最小的均方误差(mean squared errors,MSE)。最近较新的一项研究利用相似的模型,在使用共计约一千个音频片段的数据集的情况下,以与ISO12913相似的圆周模型,将音乐分类为四个象限中的一个象限,并且将音乐分类为中性、平静、快乐、悲伤、愤怒和恐惧的类别。然而,声景通常包含的不仅仅是音乐,因此仍需研究其结果是否可以泛化于更广泛的声学刺激类别。On the other hand, recent systematic reviews show that while deep neural networks are prevalent in more “objective” tasks such as sound event localization, detection, and classification, research on using deep neural networks to predict soundscape perception attribute values is scarce. In this regard, the most important research appears to be a comparison of the performance of SVM, convolutional neural networks (CNN), long short-term memory networks (LSTM), and fine-tuned VGGish networks on the Emo-Soundscapes dataset. Based on subjective pairwise comparisons using a self-assessment manikin model, the Emo-Soundscapes dataset contains 1213 audio segments ranked by valence (pleasure) and arousal (eventality). The fine-tuned VGGish model and the CNN trained from scratch exhibited the smallest mean squared errors (MSEs) in predicting the valence and arousal of the audio segments, respectively. A more recent study used a similar model, employing a circular model similar to ISO 12913, to classify music into one of four quadrants using a dataset of approximately one thousand audio clips, and further categorized music into neutral, calm, happy, sad, angry, and fearful categories. However, soundscapes typically encompass more than just music, so it remains to be investigated whether the results can generalize to a wider range of acoustic stimulus categories.

此外,文献中的这些现有模型是确定性模型,即在不考虑用作基准真实值的主观评分中固有的不确定性的情况下,相同的声学环境总是映射于相同的预测值。然而,鉴于无法合理控制的因素,例如人的当前压力级别或噪声敏感度,个人给出的这些评分本质上是随机的。因此,各个实施例可以涉及概率方法,在该概率方法中,通过训练神经网络以预测输出的可能评分的分布,并且随后从该分布中摘取,这允许模型考虑每个独特声景中不同程度的不确定性。Furthermore, these existing models in the literature are deterministic models, meaning that the same acoustic environment always maps to the same predicted value, without considering the inherent uncertainty in subjective ratings used as baseline true values. However, these ratings given by individuals are inherently random, given factors that cannot be reasonably controlled, such as a person's current stress level or noise sensitivity. Therefore, various embodiments may involve probabilistic methods in which a neural network is trained to predict a distribution of possible ratings for the output, and then extracted from that distribution, allowing the model to account for different degrees of uncertainty in each unique soundscape.

提出的方法The proposed method

图8示出根据各个实施例的自动掩蔽选择系统(automatic masker selectionsystem,AMSS)的训练和推理示意图。Figure 8 illustrates the training and inference diagrams of the automatic mask selection system (AMSS) according to various embodiments.

考虑声景和可以添加于声景的K个候选掩蔽器的集合{Mk}。以表示具有掩蔽器Mk的增强声景。换言之,是添加后的声景,其中考虑播放系统和现实世界环境的响应。找到最优的以将对增强声景的给定感知(或情感属性)f的评价最大化,相当于找到,其中f可以是例如愉悦度、事件性或其他感知属性在数字指标上的评分。然后,我们可以用增强,以获得。Consider a soundscape and a set of K candidate masks { Mk } that can be added to the soundscape. Let represent the enhanced soundscape with mask Mk . In other words, is the enhanced soundscape, where the responses of the playback system and the real-world environment are considered. Finding the optimal that maximizes the evaluation of a given perceptual (or emotional) attribute f of the enhanced soundscape is equivalent to finding , where f can be, for example, a rating on a numerical metric for pleasantness, eventiness, or other perceptual attributes. Then, we can use the enhancement to obtain .

例如,AMSS的朴素实施可以使用的某个或某种近似器,以为每个增强声景输出感知属性的预测值,并且选取具有最高的掩蔽器。然而,如前所述,该确定性输出方案忽略与感知属性指标上的主观评分相关的、固有的不确定性。例如,一个或一种声景对一些人而言可能是非常愉悦的,但对另一些人而言却可能是非常烦扰的,与此同时,另一声景对任何听众而言可能几乎普遍地是略微愉悦的,这取决于情境。For example, a naive implementation of AMSS could use one or more approximators to predict the perceptual attributes of each enhanced soundscape output and select the mask with the highest value. However, as mentioned earlier, this deterministic output scheme ignores the inherent uncertainty associated with subjective ratings on perceptual attribute metrics. For instance, one soundscape might be very pleasant to some but very disturbing to others, while another might be slightly pleasant to almost universally listeners, depending on the context.

为收集该不确定性,可以被视为具有与某个或某种“随机”声景的联合分布的随机变量,其中观察到的增强声景被视为的现实化(realization)。然后,不确定性将以的方差(variance)形式表示。在该情况下,预测模型将尝试将条件分布 建模为增强声景的某个或某种函数。为简洁起见, 被缩写为。To collect this uncertainty, it can be viewed as a random variable with a joint distribution with one or more “random” soundscapes, where the observed augmented soundscape is considered its realization. The uncertainty is then expressed as the variance of . In this case, the predictive model will attempt to model the conditional distribution as one or more functions of the augmented soundscape. For simplicity, is abbreviated as .

神经近似器806可以用于输出。神经近似器806可以称为概率感知属性预测器(probabilistic perceptual attribute predictor,PPAP)。在图8所示的、所提出的AMSS中,每个掩蔽器均可传递至前述PPAP 806,PPAP 806输出某个或某种预定义分布族的一个或多个参数。PPAP 806可以被训练以输出正态分布的平均值和对数标准差,平均值和对数标准差被用于对输出 进行建模。一旦计算出足够的(或全部的)预测分布,便可基于某个或某种预定标准选择“最佳”掩蔽器。例如,可以仅基于最高的选择掩蔽器,可从中采样相似于或类似于的确定值的集合以激励掩蔽器探索,或者可以使用考虑的、较复杂的标准。A neural approximator 806 can be used for the output. The neural approximator 806 can be called a probabilistic perceptual attribute predictor (PPAP). In the proposed AMSS shown in Figure 8, each mask can be passed to the aforementioned PPAP 806, which outputs one or more parameters of one or more predefined distribution families. The PPAP 806 can be trained to output the mean and log-standard deviation of a normal distribution, which are used to model the output. Once sufficient (or all) predictive distributions have been computed, the “best” mask can be selected based on one or more predetermined criteria. For example, the mask can be selected based solely on the highest value, a set of similar or analogous values can be sampled to incentivize mask exploration, or more complex criteria can be considered.

尽管只有的观测值是可用的,并且的真实分布是不明的,但是仍然可以在给定输出分布的情况下,通过将基准真实值(ground truth)的对数概率以受贝叶斯优化(Bayesian optimization)启发的方式最大化,从而优化模型。在给定声景、掩蔽器的集合和基准真实值的情况下,对声景的损失函数的贡献可以由下式给出:Although only a few observations are available and the true distribution is unknown, the model can still be optimized given the output distribution by maximizing the log probability of the ground truth in a Bayesian optimization-inspired manner. Given the soundscape, the set of masks, and the ground truth, the contribution of the loss function to the soundscape can be given by the following equation:

(1)(1)

,(2)(2)

其中是输出分布的对数密度函数,其中省略加性常数。在训练过程中,可以通过具有可用基准真实值的、批量的声景-掩蔽器对,以优化模型。Here, is the logarithmic density function of the output distribution, where the additive constant is omitted. During training, the model can be optimized using batches of soundscape-mask pairs with available benchmark true values.

如计算式(2)所示,使用正态分布的损失函数可以被认为是由对数标准差正则化(regularized)的加权MSE的损失。该损失函数相对于是固有地稳定的,因为第一项激励(encourage)较大的,与此同时,第二项激励较小的。当然,输出分布的一些其他选择可以类似地缩减(reduce)为其他偏差指标(deviation measures),例如缩减为正则化的平均绝对误差的拉普拉斯分布(Laplace distribution)。应注意的是,也可以通过将σk设置为某个预定常数且使用基准(ground)与之间的、纯MSE的损失训练模型,以强制模型进行确定性学习。为了进行消融研究(ablation study)的目的,可以使用下式As shown in Equation (2), the loss function using a normal distribution can be considered as a loss of weighted MSE regularized by log-standard deviation. This loss function is inherently stable relative to the first incentive term, while the second incentive term is smaller. Of course, some other choices of the output distribution can be similarly reduced to other deviation measures, such as a regularized Laplace distribution of the mean absolute error. It should be noted that deterministic learning can also be forced by setting σk to some predetermined constant and training the model using a loss of pure MSE between the ground and the ground. For the purpose of ablation studies, the following equation can be used.

(3)(3)

作为计算式(2)的确定性对应项。计算式(3)中的确定性损失可以被认为是具有静态=1的计算式(2)。As the deterministic counterpart of formula (2), the deterministic loss in formula (3) can be considered as formula (2) with static = 1.

验证实验Verification Experiment

为验证所提出的系统,我们令为ISO12913-3标准定义的归一化的愉悦度量度,也可以称为ISO愉悦度。具体而言,To validate the proposed system, we use the normalized pleasure metric defined by the ISO 12913-3 standard, also known as the ISO pleasure metric. Specifically,

(4)(4)

其中分别是以5点李克特量表表示的、参与者认为增强声景愉悦、烦扰、平静、混乱、活跃和单调的程度。对于每个,可以预测分布,其中“基准真实值”标签是参与者给出的、对的ISO愉悦度评分。作为消融研究,将使用所提出的方法所训练的模型与通过直接预测的确定性模型进行比较。验证实验是在对各个增强声景的主观反应的数据集上进行的。These represent the degree to which participants perceived the enhanced soundscape as pleasant, disturbing, calm, chaotic, active, and monotonous, expressed using a 5-point Likert scale. For each, a predictive distribution can be generated, where the “benchmark true value” label is the participant’s ISO pleasantness rating. As an ablation study, the model trained using the proposed method will be compared with a deterministic model based on direct prediction. Validation experiments were conducted on a dataset of subjective responses to each enhanced soundscape.

数据集Dataset

该数据集包含5折交叉验证集(5-fold cross-validation set)中的、对增强声景的主观反应的、12600个去重集合(unique sets)以及独立测试集(independent test set)中的反应的48个附加集合,总共12648个样本。在给定对ISO12913-2情感反应问卷的以下6项子集的反应的集合:“您在多大程度上同意或反对当前周围声音环境是{愉悦的、混乱的、活跃的、平静的、烦扰的、单调的}?”的情况下,每个样本将增强声景(作为原始音频记录)映射于ISO愉悦度值。This dataset contains 12,648 samples in total, including 12,600 unique sets of subjective responses to the enhanced soundscape from the 5-fold cross-validation set, and 48 additional sets of responses from the independent test set. Given a set of responses to the following six subsets of the ISO 12913-2 Affective Response Questionnaire: “To what extent do you agree or disagree that your current surrounding sound environment is {pleasant, chaotic, active, calm, disturbing, monotonous}?”, each sample maps the enhanced soundscape (as a raw audio recording) to an ISO pleasantness value.

参与者按照5分量表进行回答,其中该5分量表具有多个标签(labels)“非常反对”、“反对”、“既不同意也不反对”、“同意”和“非常同意”,该多个标签分别编码为1、2、3、4和5。然后,将编码后的回答用于计算式(4)中,以为每个增强声景计算ISO愉悦度的基准真实值标签。Participants responded using a 5-point scale with multiple labels: “Strongly Disagree,” “Disagree,” “Neither Agree nor Disagree,” “Agree,” and “Strongly Agree,” coded as 1, 2, 3, 4, and 5, respectively. The coded responses were then used in Equation (4) to calculate the baseline true value label for ISO pleasure level for each enhanced soundscape.

增强声景Enhance soundscape

5折交叉验证集中的增强声景是通过将Freesound和xeno-canto的录音的30秒摘录作为“掩蔽器”添加于世界城市声景(Urban Soundscapes of the World,USotW)数据库的声景的立体声录音(binaural recordings)的30秒摘录以制成的,世界城市声景数据库是全面的城市声景数据集。也包括没被增强的声景以作为对照。所使用的全部录音均以44.1kHz采样。The enhanced soundscapes in the 5-fold cross-validation set were created by adding 30-second extracts of stereo recordings of soundscapes from the Urban Soundscapes of the World (USotW) database as “masks” to the original recordings. The USotW database is a comprehensive dataset of urban soundscapes. Unenhanced soundscapes were also included as a control. All recordings used were sampled at 44.1 kHz.

每一折具有56个掩蔽器的组,该56个掩蔽器的组分为以下类别:鸟(16)、施工(8)、交通(8)、水(16)和风(8)。选择这些类别以涵盖总体上被评估为愉悦的和烦扰的声音类型范围。每一折中的掩蔽器和声景都是互斥的(disjoint)。Each fold contains a group of 56 maskers, which are categorized as follows: birds (16), construction (8), traffic (8), water (16), and wind (8). These categories were chosen to cover a range of sound types that were generally assessed as pleasant and disturbing. The maskers and soundscapes in each fold are disjoint.

测试集中的增强声景以相似的方式制成,其中7个掩蔽器(包括没被增强的对照掩蔽器在内共8个掩蔽器)独立于交叉验证集中的掩蔽器,并且声景的6个立体声录音独立于USotW数据集,该6个立体声录音使用相同的声景索引协议(Soundscape IndicesProtocol)录制。7个掩蔽器摘录自xeno-canto音轨ID(识别号)640568(鸟“B1”)和568124(鸟“B2”),以及Freesound音轨ID586168(施工“Co”)、587219(交通“Tr”)、587000(水“W1”)、587759(水“W2”)和587205(风“Wi”)。掩蔽器被完全地(exhaustively)添加于测试集的全部声景,这在测试集中产生48个样本。在为测试集以恒定的、0dB的声景与掩蔽器的比率以及为交叉验证集以dB为单位、从{-6,-3,0,3,6}中随机选定的值添加掩蔽器之前,全部声景均根据K.Ooi等人描述的方法(“在人造头上进行立体声耳机音频校准的自动化”,MethodsX,第8卷,第2期,第101288页,2021年)校准为实时测量的A加权等效的SPL(。The enhanced soundscapes in the test set were created in a similar manner, with seven masks (eight masks in total, including the unenhanced control mask) independent of the masks in the cross-validation set, and six stereo recordings of the soundscapes independent of the USotW dataset, recorded using the same Soundscape Indices Protocol. The seven masks were extracted from xeno-canto track IDs 640568 (bird "B1") and 568124 (bird "B2"), and Freesound track IDs 586168 (construction "Co"), 587219 (traffic "Tr"), 587000 (water "W1"), 587759 (water "W2"), and 587205 (wind "Wi"). The masks were exhaustively added to all soundscapes in the test set, resulting in 48 samples. Before adding a mask to the test set with a constant, 0 dB soundscape to mask ratio and to the cross-validation set with a mask in dB values randomly selected from {-6,-3,0,3,6}, all soundscapes were calibrated to an A-weighted equivalent SPL measured in real time according to the method described by K. Ooi et al. (“Automation of Stereo Headphone Audio Calibration on an Artificial Head”, MethodsX, Vol. 8, No. 2, p. 101288, 2021).

主观反应Subjective reaction

为获得对增强声景的主观反应,招募300名参与者,每人对从验证集的一个折中随机选择的42个独特的增强声景进行评分。招募另外5名参与者,每人对测试集中的48个增强声景进行评分。因此,验证集和测试集中的每个增强声景分别由一名参与者和五名参与者评分。所有参与者都使用一副由外部声卡(Creative Sound Blaster E5)驱动的耳罩式耳机(Beyerdynamic Custom One Pro)以收听经校准的增强声景。听完每个增强声景后,他们回答前文描述的6项问卷(在“数据集”标题下),并且根据计算式4从他们的回答中计算出ISO愉悦度。To obtain subjective responses to the enhanced soundscapes, 300 participants were recruited, each rating 42 unique enhanced soundscapes randomly selected from a compromise of the validation set. An additional 5 participants were recruited, each rating 48 enhanced soundscapes in the test set. Thus, each enhanced soundscape in the validation and test sets was rated by one participant and five participants, respectively. All participants listened to the calibrated enhanced soundscapes using a pair of over-ear headphones (Beyerdynamic Custom One Pro) driven by an external sound card (Creative Sound Blaster E5). After listening to each enhanced soundscape, they answered the 6 questionnaires described above (under the heading “Datasets”), and ISO pleasantness was calculated from their responses according to Equation 4.

模型与训练Model and Training

增强声景的对数梅尔频谱图(Log-mel spectrograms)被用作输入。使用具有50%重叠率的4096样本汉恩窗口(Hann window)提取对数梅尔频谱图,并且将该对数梅尔频谱图压缩至64个梅尔箱(mel bins)。图9示出(a)根据各个实施例使用的基本卷积循环神经网络(CRNN)架构;(b)根据各个实施例,(a)中的特征映射块的一个可能实施;以及(c)根据各个其他实施例,(a)中的特征映射块的其他可能实施。研究四个不同的特征映射块,即使用双向门控循环单元(bidirectional gated recurrent unit,BiGRU)的基础映射块(vanilla mapping block)、加性注意力块(additive attention block)、点积注意力块(dot-product attention block)和具有4个头的多头注意力块。基础映射块如图9(b)所示,与此同时,共享相同通用工作流程的、全部的基于注意力的块如图9(c)所示。图9(a)的模型的最后数层是全连接层(dense layers),该全连接层最终输出μk和。在确定性消融模型中,被忽略。Log-mel spectrograms for enhanced soundscapes were used as input. Log-mel spectrograms were extracted using a 4096-sample Hann window with 50% overlap, and then compressed to 64 mel bins. Figure 9 illustrates (a) the basic convolutional recurrent neural network (CRNN) architecture used according to various embodiments; (b) a possible implementation of the feature mapping block in (a) according to various embodiments; and (c) other possible implementations of the feature mapping block in (a) according to various other embodiments. Four different feature mapping blocks were investigated: a vanilla mapping block using a bidirectional gated recurrent unit (BiGRU), an additive attention block, a dot-product attention block, and a multi-head attention block with four heads. The vanilla mapping block is shown in Figure 9(b), while all attention-based blocks sharing the same general workflow are shown in Figure 9(c). The last few layers of the model in Figure 9(a) are dense fully connected layers, which ultimately output μk and . In the deterministic ablation model, these are ignored.

全部模型均使用5折交叉验证方案进行训练,其中每一折具有10个模型,总共50个模型。每一折为多个模型使用相同的10个种子。全部模型均使用亚当优化器(Adamoptimizer)以5×10−5的学习率训练最多100轮。对于每个模型,在交叉验证集和测试集二者中均使用具有最佳验证损失的模型权重进行评估。All models were trained using a 5-fold cross-validation scheme, with 10 models per fold, for a total of 50 models. Each fold used the same 10 seeds for multiple models. All models were trained using the Adam optimizer with a learning rate of 5 × 10⁻⁵ for a maximum of 100 epochs. For each model, the model weights with the optimal validation loss were evaluated on both the cross-validation and test sets.

结果与讨论Results and Discussion

图10是示出根据各个实施例的概率感知属性预测器(PPAP)在针对每个设置测试的10次运行中的均折均方误差(mean fold mean squared errors,MSEs)(±标准差)(上:交叉验证集,下:测试集)。每次运行可以包括五个模型,每个模型在交叉验证集的不同折上,但是具有相同的初始条件。对于全部模型,对MSE的贡献被计算为。星号(*)表示统计上显著的改进(p<0.05)。图10总结前文描述的验证实验在5折交叉验证集和独立测试集二者上的结果。作为参考,还提供一个简单的(trivial)确定性“标签均值模型”的结果。在该“标签均值模型”中,训练集中的标签均值用作验证集和测试集中全部刺激(stimuli)的预测。所有其他研究模型的表现均优于标签均值模型,因此表明由其他研究模型进行的特征提取是有益的。这可能是因为,与验证集(12600个样本、280个掩蔽器)相比,测试集较小且多样性较低(48个样本、7个掩蔽器)。Figure 10 illustrates the mean fold mean squared errors (MSEs) (± standard deviation) of the Probabilistic Aware Attribute Predictor (PPAP) across 10 runs for each test setting according to various embodiments (top: cross-validation set, bottom: test set). Each run may include five models, each on a different fold of the cross-validation set, but with the same initial conditions. For all models, the contribution to the MSE is calculated. An asterisk (*) indicates a statistically significant improvement (p<0.05). Figure 10 summarizes the results of the validation experiments described above on both the 5-fold cross-validation set and the independent test set. For reference, results for a simple (trivial) deterministic “label mean model” are also provided. In this “label mean model”, the label mean from the training set is used as the prediction for all stimuli in the validation and test sets. All other research models outperform the label mean model, thus demonstrating the benefit of feature extraction by other research models. This may be because the test set is smaller and less diverse (48 samples, 7 masks) compared to the validation set (12,600 samples, 280 masks).

此外,对于所测试的四个架构,具有概率输出的模型表现优于具有验证集和测试集二者的确定性输出的等同模型。图10还示出每个模型的MSE降低百分比。可以观察到,具有多头注意力块的CRNN和基础CRNN在验证集和测试集的MSE中分别呈现1.1%和7.8%的最大改进。Furthermore, for the four architectures tested, the model with probabilistic output outperformed the equivalent model with deterministic output on both the validation and test sets. Figure 10 also shows the percentage reduction in MSE for each model. It can be observed that the CRNN with multi-head attention blocks and the basic CRNN showed the largest improvements in MSE on the validation and test sets, at 1.1% and 7.8%, respectively.

为使这些降低的显著性(significance)量化,对验证集MSE和测试集MSE二者进行确定性模型与概率模型之间的双侧威尔科克森符号秩测试(two-sided Wilcoxon signed-rank tests)。在图10中,在5%显著性级别上显著的降低以星号标记。对于基础CRNN和具有加性注意力的PPAP,测试集MSE的降低在统计上是显著的,并且提供证据以支持所提出的方法对没有看到的数据的泛化性(generalizability)。这可以归因于测试集参与者贡献的固有随机性与训练集中的参与者不同,并且PPAP对该随机性的泛化较好。然而,需要进一步的消融研究以验证该假设。To quantify the significance of these reductions, two-sided Wilcoxon signed-rank tests were performed on both the validation and test set MSEs, comparing deterministic and probabilistic models. In Figure 10, reductions significant at the 5% significance level are marked with an asterisk. For both the base CRNN and PPAP with additive attention, the reduction in test set MSE was statistically significant, providing evidence to support the generalizability of the proposed method on unseen data. This can be attributed to the inherent randomness of the contributions from test set participants, which differs from that of participants in the training set, and PPAP's good generalization to this randomness. However, further ablation studies are needed to validate this hypothesis.

然而,测试集MSE的最低改进实际上是在具有多头注意力块的PPAP中观察到的。这可以归因于具有多头注意力块的PPAP比其他三个模型具有多出约50%的参数(约120K对约80K)的事实,这可能导致训练模型(确定性训练模型和概率性训练模型二者)过度拟合于相对较小的12600个训练和验证样本的数据集。However, the lowest improvement in test set MSE was actually observed in PPAP with multi-head attention blocks. This can be attributed to the fact that PPAP with multi-head attention blocks has about 50% more parameters than the other three models (about 120K vs. about 80K), which may cause the trained models (both deterministic and probabilistic) to overfit the relatively small dataset of 12,600 training and validation samples.

图11示出使用测试集上的50个模型(点积变量)进行掩蔽选择,其中采用根据各个实施例的朴素最大选择方案(左)以及根据各个实施例的随机采样方案以激励掩蔽器探索(右)。在每个子图的顶部上示出用于每个基本声景的选择标准。单元格颜色表示测试集参与者评定的平均ISO愉悦度。圆圈尺寸表示选择该掩蔽器的模型数量。掩蔽器描述已在前文提供(在标题“增强声景”下)。Figure 11 illustrates mask selection using 50 models (dot product variables) on the test set, employing a naive maximum selection scheme (left) and a random sampling scheme (right) to incentivize masker exploration, according to various embodiments. The selection criteria used for each basic soundscape are shown at the top of each subplot. Cell color represents the average ISO pleasure rating by test set participants. Circle size indicates the number of models selected for that mask. Masker descriptions have been provided previously (under the heading "Enhanced Soundscapes").

在朴素选择方案中,大多数模型在全部基本声景中选择第一鸟鸣掩蔽器(B1)。这与先前声景研究中的共识一致,该先前声景研究总体上观察到对鸟鸣掩蔽器的、较高愉悦度的反应。然而,第二鸟鸣掩蔽器(B2)可能总体上获得参与者的较低评分,并且相应地,模型的投票数较低,这是因为它在测试集声景的情境中被认为愉悦度较低,这在先前的研究中已经观察到。In the naive choice scenario, most models selected the first birdcall masker (B1) across all basic soundscapes. This aligns with the consensus in previous soundscape studies, which generally observed higher pleasantness responses to birdcall maskers. However, the second birdcall masker (B2) likely received lower overall participant ratings, and consequently, fewer model votes, because it was perceived as less pleasant in the context of the test set soundscape, as observed in previous studies.

在随机选择方案中,系统仍然倾向于选择为大多数声景提供良好的愉悦度得分的B1,但是现在,系统对其他掩蔽器的探索比朴素方案多得多。这在实时系统中可以是有益的,在实时系统中可以实时地获得人员反馈,以自适应地调整掩蔽选择或改进未来的模型。还可以看出,在声景5的情况下,没被增强的声景的愉悦度已经相对较高,探索率高于其他基本声景,这允许在愉悦度级别上仅做出较小妥协的情况下,提供较多样化的声学体验。In the random selection scenario, the system still tends to choose B1, which provides a good pleasure score for most soundscapes. However, the system now explores other masks much more extensively than in the naive scenario. This can be beneficial in real-time systems where human feedback can be obtained in real time to adaptively adjust mask selection or improve future models. It can also be seen that in the case of soundscape 5, the unenhanced soundscape already has a relatively high pleasure score and a higher exploration rate than other basic soundscapes, allowing for a more diverse acoustic experience with only minor compromises in pleasure level.

结论in conclusion

各个实施例可以涉及用于以人为本的城市声景增强的自动掩蔽选择系统(AMSS),该系统使用概率感知属性预测器(PPAP)。所提出的PPAP是使用卷积循环神经网络实施的,并且经过训练从而以概率方式输出预测。这允许该系统在考虑声学刺激的人类主观感知中固有的随机性的同时,预测声景的感知属性。通过一项有超过300名参与者和超过12000个独特声景的大规模收听测试,我们验证了我们的PPAP在预测增强声景的愉悦度方面的有效性,该增强声景包括从没被看到的声景和掩蔽器中生成的声景。AMSS的未来工作可能包括实时实施以评估其生态效度(ecological validity),以及对所提出的方法对例如事件性或平静度的其他感知属性(f)的研究,这是因为所提出的方法不是专用于任何特定属性的。事实上,由于PPAP所依据的主要假设是随机的基准真实值标签的主要假设,人们也可以设想将其应用于任何需要预测主观评价的情况。Various embodiments may relate to an Automatic Masking Selection System (AMSS) for human-centered urban soundscape enhancement, which uses a Probabilistic Perceived Attribute Predictor (PPAP). The proposed PPAP is implemented using a convolutional recurrent neural network and trained to output predictions probabilistically. This allows the system to predict the perceptual attributes of soundscapes while taking into account the inherent randomness in the human subjective perception of acoustic stimuli. Through a large-scale listening test with over 300 participants and over 12,000 unique soundscapes, we validated the effectiveness of our PPAP in predicting the pleasantness of enhanced soundscapes, which included soundscapes generated from unseen and masked soundscapes. Future work on AMSS may include real-time implementation to assess its ecological validity, and studies of the proposed method on other perceptual attributes (f), such as eventiness or calmness, since the proposed method is not specific to any particular attribute. In fact, since the main assumption upon which PPAP is based is the main assumption of randomized baseline true value labeling, it can also be envisioned for application to any situation requiring the prediction of subjective evaluations.

涉及对给定声景添加称为“掩蔽器”的声音的声景增强是以人为本的城市噪声缓解措施,该措施旨在提高整体声景质量。然而,掩蔽器的选择通常是基于繁琐的过程预测的,并且对于现实世界声景的时变性而言是不灵活的。由于每个声景的感知独特性以及人类感知固有的主观性,本申请提出了概率感知属性预测器(PPAP),该概率感知属性预测器预测随机分布的参数作为输出,以代替单个确定值作为输出。在使用PPAP的情况下,本申请开发了自动掩蔽选择系统(AMSS),该系统基于用于给定声景的ISO12913-3愉悦度评分的预测分布,选择最佳掩蔽器候选。通过有300名参与者的大规模收听测试,收集了12600个主观反应,每个主观反应针对一个独特的增强声景,从而以5折交叉验证方案训练PPAP模型。在使用卷积循环神经网络主干以及具有为PPAP进行的注意力机制的数个变量的测试的情况下,本申请使用具有48个没被看到的增强声景的盲测集,进行了对所提出的系统的评估,以评价概率输出方案相对于常规确定性系统的有效性。Soundscape enhancement, which involves adding sounds called “masks” to a given soundscape, is a human-centered urban noise mitigation measure aimed at improving the overall soundscape quality. However, mask selection is often based on cumbersome predictive processes and is inflexible in the face of the time-varying nature of real-world soundscapes. Due to the perceptual uniqueness of each soundscape and the inherent subjectivity of human perception, this application proposes a Probabilistic Perceptual Attribute Predictor (PPAP), which predicts randomly distributed parameters as outputs instead of a single deterministic value. Using PPAP, this application develops an Automatic Mask Selection System (AMSS) that selects the best mask candidate based on a predicted distribution of ISO 12913-3 pleasantness ratings for a given soundscape. A PPAP model was trained using a 5-fold cross-validation scheme by collecting 12,600 subjective responses, each for a unique enhanced soundscape, through a large-scale listening test with 300 participants. In tests using a convolutional recurrent neural network backbone and several variables with an attention mechanism for PPAP, this application evaluated the proposed system using a blind test set of 48 unseen enhanced soundscapes to assess the effectiveness of the probabilistic output scheme relative to conventional deterministic systems.

示例研究2Example Study 2

介绍introduce

城市噪声污染的减弱是复杂且多方面的问题,其中噪声源的消除或减弱通常不是切实可行的选择。与旨在降低声压级或声能的常规噪声控制方法不同,声景方法提出基于感知声学舒适度的、较全面的、以人为本的策略。这种通常称为声景增强的技术涉及将“所期望出现的”声音添加于声学环境中,以“掩蔽”噪声且改善整体感知声学质量,这种技术在虚拟环境和现实环境二者中,针对各个类型的噪声源都取得了令人满意的结果。Reducing urban noise pollution is a complex and multifaceted problem, where eliminating or reducing noise sources is often not a practical option. Unlike conventional noise control methods that aim to reduce sound pressure levels or sound energy, the soundscape approach proposes a more comprehensive, human-centered strategy based on perceived acoustic comfort. This technique, often called soundscape enhancement, involves adding "desired" sounds to the acoustic environment to "mask" noise and improve overall perceived acoustic quality. This technique has achieved satisfactory results for various types of noise sources in both virtual and real environments.

然而,“掩蔽器”和播放级别的选择常规上是任意的、专家指导的,或者是基于事后分析的。这些掩蔽选择过程不仅通常是耗时耗力的,而且对于现实世界声景的动态特性而言是不灵活的。因此,这些掩蔽器和播放级别的静态选择通常不可能始终实现最佳的声学舒适度。为解决这些限制,示例研究1示出通过由概率深度学习模型驱动的掩蔽选择系统的使用,实现的适应性声景增强的一个首次尝试,该概率深度学习模型为特定增强声景预测“愉悦度”的概率分布。However, the selection of “masks” and playback levels is typically arbitrary, expert-guided, or based on post-hoc analysis. These masking selection processes are not only often time-consuming and labor-intensive, but also inflexible regarding the dynamic characteristics of real-world soundscapes. Therefore, these static selections of masks and playback levels often cannot consistently achieve optimal acoustic comfort. To address these limitations, Example Study 1 demonstrates a first attempt at adaptive soundscape enhancement using a masking selection system driven by a probabilistic deep learning model that predicts a probability distribution of “pleasure” for a specific enhanced soundscape.

尽管上述概率感知属性预测器(PPAP)模型为深度学习驱动的自动声景增强提供令人信服的概念证明,但是该模型可能不是为实时部署设计的。该PPAP模型的实施需要将增强声景的预混音轨作为输入。混合过程是在相对于所录制的没被增强的声景音轨的已知的实时SPL级别、以预先校准的声景与掩蔽器的比率完成的。这可能会导致无视掩蔽增益级别的模型--掩蔽增益级别是实时部署的关键播放参数。此外,要求将基本声景记录与掩蔽器的混合作为输入,给整个系统增加不必要的复杂性。首先,全部候选掩蔽器都需要播放所需的、处于特定声景与掩蔽器的比率(SMR)的数字增益查对表。其次,在选择过程中,每个候选掩蔽器-SMR对都需要计算增强声景音轨,这增加了系统的计算成本(computationaloverhead)。While the Probabilistic Attribute Predictor (PPAP) model described above provides a compelling proof of concept for deep learning-driven automatic soundscape enhancement, it may not be designed for real-time deployment. Implementing this PPAP model requires a premixed track of the enhanced soundscape as input. The mixing process is performed at a pre-calibrated soundscape-to-mask ratio relative to a known real-time SPL level on the recorded, unenhanced soundscape track. This could lead to a model that ignores the masking gain level—a critical playback parameter for real-time deployment. Furthermore, requiring the mixing of the base soundscape recording with the mask as input adds unnecessary complexity to the entire system. First, all candidate masks require a digital gain lookup table that plays the desired soundscape-to-mask ratio (SMR) for a given soundscape. Second, during the selection process, each candidate mask-SMR pair requires computation of the enhanced soundscape track, increasing the system's computational overhead.

各个实施例可以呈现具有修改的计算式的PPAP模型,该PPAP模型将允许在现实世界的部署中实现较高效的计算和带宽系统。所提出的模型可以在输入阶段将基本声景、掩蔽器和增益级别相互解耦,并且替代地为基本声景和掩蔽器引入单独的特征提取器分支。通过使用掩蔽器的、独立于实时声压级(SPL)的数字增益级别作为调节性输入,“增强”过程仅发生于特征空间以代替波形空间。然后基于基本声景特征、掩蔽特征和增益调节特征,使用注意力机制以“查询”目标感知属性的分布。所提出的方法导致系统多个组件的复杂性降低。首先,可以不需要将声景的原始波形发送至推理引擎以与掩蔽器混合,因为可以使用带宽消耗较少的频谱数据作为模型输入。对于云推理,这意味着边缘的数据流出率(dataegress rate)的显著降低。其次,掩蔽特征可以独立于增益级别进行预计算,这减少了推理时间,并且因此减少了总体延迟(latency)。第三,现在可以在无需重新计算增强声景或甚至声景和掩蔽特征的情况下,较快地对同一掩蔽器的多个增益级别进行属性预测。最后,现在可以在无需创建查对表的情况下,实现新掩蔽器对系统的添加。Various embodiments can present a PPAP model with a modified computational formula, which will allow for more efficient computational and bandwidth systems in real-world deployments. The proposed model can decouple the base soundscape, mask, and gain level from each other at the input stage, and instead introduce separate feature extractor branches for the base soundscape and mask. The "enhancement" process occurs only in the feature space instead of the waveform space, using the mask's digital gain level, independent of the real-time sound pressure level (SPL), as a modulating input. Then, based on the base soundscape features, masking features, and gain modulating features, an attention mechanism is used to "query" the distribution of the target's perceptual attributes. The proposed method leads to a reduction in the complexity of several system components. First, it is not necessary to send the raw waveform of the soundscape to the inference engine for mixing with the mask, as less bandwidth-intensive spectral data can be used as model input. For cloud inference, this means a significant reduction in the data egress rate at the edge. Second, the masking features can be pre-computed independently of the gain level, which reduces inference time and thus overall latency. Third, attribute prediction for multiple gain levels of the same mask can now be performed faster without recalculating the enhanced soundscape or even the soundscape and masking features. Finally, new masks can now be added to the system without creating lookup tables.

提出的方法The proposed method

准备工作Preparation

考虑数字化声景记录 ,其中C≥1是通道数,并且是样本索引。通过表示si的、以dBA为单位的实时A加权等效声压级。为简洁起见,所有提到的声压级(SPL)均指30秒,除非另有说明。假设所有声景均由相同的实时记录设置记录,并且在假设近似线性的情况下,声压与数字级别的比率(sound pressure-to-digitallevel ratio,SPDR),即线性尺度上的相对声压级与设置的数字满量程(digital fullscale,DFS)的比率,可以认为大致恒定,并且以表示,其单位是,其中是基准声压级。由于全部声景输入都源自同一录音设备,这也意味着实时SPL信息隐含地嵌入输入频谱图数据中。Consider a digital soundscape recording, where C≥1 is the number of channels and is the sample index. The real-time A-weighted equivalent sound pressure level (SPL) of s<sub> i </sub> is expressed in dBA. For simplicity, all SPL references refer to 30 seconds unless otherwise stated. It is assumed that all soundscapes are recorded using the same real-time recording setup, and that, under the assumption of approximate linearity, the sound pressure-to-digital-level ratio (SPDR), i.e., the ratio of the relative SPL on a linear scale to the set digital full-scale (DFS), can be considered approximately constant, and is denoted by , with units of , where is the reference SPL. Since all soundscape inputs originate from the same recording device, this also means that real-time SPL information is implicitly embedded in the input spectrogram data.

考虑单通道掩蔽器 。对于掩蔽器,假设它们不都是由同一录音装置录制的。在实践中,这是因为掩蔽器通常源自与实时录音装置不同的录音系统,并且通常通过例如Freesound和Xeno-canto的开放内容供应商获得。因此,每个掩蔽器的SPDR通常是不明的(unknown),并且不能假设其与任何其他掩蔽器的SPDR相同。以表示掩蔽器的SPDR。在(通过对本文中使用的全部掩蔽器进行彻底的增益校准以验证的)近似线性的假设下,可以通过将掩蔽器以定比例,以将掩蔽器归一化为与声景录音相同的SPDR。如果和实时SPL 二者是已知的,则在无需进一步校准的情况下,具有特定声景与掩蔽器的比率(SMR)的、以dBA为单位的声景增强可以相对容易地进行。然而,在实践中通常是不明的,因此常规解决方法是在使用具有的数字级别与声压的比率(digital level-to-sound pressure ratio,DSPR)的播放系统的情况下,进行校准,以在的多个值中,获得以的声压级再现所需的数字增益。获得 的精确值所需的校准通常可能需要专门的设备,例如具有人造头的录音装置,并且可能是耗时的和/或耗力的,这取决于所使用的校准方法。在本文中,明确用作调节性输入,因此去除了SPL相关的SMR的需求或掩蔽校准的需求。Consider single-channel masks. For masks, assume they are not all recorded by the same recording device. In practice, this is because masks are often derived from recording systems different from the real-time recording device and are typically obtained through open content providers such as Freesound and Xeno-canto. Therefore, the SPDR of each mask is usually unknown and cannot be assumed to be the same as the SPDR of any other mask. Let represent the SPDR of the mask. Under the assumption of approximate linearity (verified by thorough gain calibration of all masks used in this paper), the mask can be normalized to the same SPDR as the soundscape recording by scaling the mask. If both the soundscape and the real-time SPL are known, soundscape enhancement in dBA units with a specific soundscape-to-mask ratio (SMR) can be performed relatively easily without further calibration. However, in practice, this is often unclear, so the conventional solution is to calibrate using a playback system with a digital level-to-sound pressure ratio (DSPR) to obtain the digital gain required for sound pressure level reproduction from among multiple values of . The calibration required to obtain an accurate value of often requires specialized equipment, such as a recording device with an artificial head, and can be time-consuming and/or labor-intensive, depending on the calibration method used. In this paper, is explicitly used as an adjustable input, thus eliminating the need for SMR or masking calibration related to SPL.

模型Model

图12A示出根据各个实施例的概率感知属性预测器(PPAP)的示意图。设为声景的对数梅尔频谱图,并且是掩蔽器的对数梅尔频谱图,其中是频谱图时间帧的数量,并且是梅尔箱的数量。在本文中,,,并且对应30秒的信号,该信号以44.1kHz、以4096个样本的短时傅里叶变换(short-time Fourier transform,STFT)窗口尺寸以及50%的帧重叠率采样。声景和掩蔽嵌入分别由特征提取器1222 和特征提取器1226 计算,以使得Figure 12A illustrates a schematic diagram of the Probabilistic Aware Attribute Predictor (PPAP) according to various embodiments. Let be the log-Mel spectrogram of the soundscape, and be the log-Mel spectrogram of the mask, where is the number of time frames of the spectrogram, and is the number of Mel bins. In this paper, , , and correspond to a 30-second signal sampled at 44.1 kHz with a short-time Fourier transform (STFT) window size of 4096 samples and a 50% frame overlap rate. The soundscape and mask embeddings are computed by feature extractors 1222 and 1226, respectively, such that

,(5)(5)

(6)(6)

其中是嵌入维度,并且是特征时间帧的数量。每个特征提取器1222、1226可以包括5个卷积块,每个卷积块包含具有3×3的核和1步幅的卷积层;批量归一化;具有概率0.2的退出率(dropout);Swish激活;以及2×2的平均池化层。卷积层按顺序包含16、32、48、64和64个输出通道。Where is the embedding dimension and the number of feature time frames. Each feature extractor 1222, 1226 can include 5 convolutional blocks, each containing a convolutional layer with a 3×3 kernel and a stride of 1; batch normalization; dropout with a probability of 0.2; Swish activation; and a 2×2 average pooling layer. The convolutional layers contain 16, 32, 48, 64, and 64 output channels in sequence.

在使用注意力的查询-键-值(query-key-value,QKV)模型的情况下,感知属性预测过程可以看作是使用掩蔽器查询映射,其中没被增强的声景给出键,并且增强声景给出值。我们考虑在嵌入域中执行“增强”(在特征增强块1230中),以代替在音频或频谱域中执行声景增强。这可以通过使用以掩蔽数字增益调节的简单映射完成,以使得In the case of an attention-based query-key-value (QKV) model, the perceptual attribute prediction process can be viewed as using a masked query map, where the unenhanced soundscape provides the key, and the enhanced soundscape provides the value. We consider performing "enhancement" in the embedding domain (in feature enhancement block 1230) instead of performing soundscape enhancement in the audio or spectral domain. This can be accomplished using a simple map adjusted to mask digital gain, so that...

,(7)(7)

其中是查询增益级别,,是增益调节增强层。有的三个不同的实施,即This refers to the query gain level, which is the gain adjustment enhancement layer. There are three different implementations, namely...

(8)(8)

(9)(9)

(10)(10)

其中张量|是最后一个轴上的张量拼接,‖是创建附加轴的张量堆叠运算符,是具有D个输出单元的全连接层,是具有D个过滤器的、使用一维核将堆叠的维度压缩为单一轴的卷积层。图12B示出根据各个实施例中的计算式(8)、(9)和(10)的三种特征增强方法。在训练过程中,如果掩蔽器是静音音轨(即无增强),我们将使随机化,其中和是具有非静音掩蔽器的训练样本的对数增益的平均值(mean)和标准差(standard deviation)。这是为了教导模型忽略静音掩蔽器的增益值。在查询、键和值嵌入到位的情况下,可以使用任何QKV兼容的注意力块1220a ,计算输出嵌入,以使得Where tensor | is the tensor concatenation on the last axis, || is the tensor stacking operator that creates the additional axis, is a fully connected layer with D output units, and is a convolutional layer with D filters that uses a one-dimensional kernel to compress the stacked dimensions into a single axis. Figure 12B illustrates three feature enhancement methods according to the computation of equations (8), (9), and (10) in the various embodiments. During training, if the mask is a silent track (i.e., no enhancement), we will randomize, where and are the mean and standard deviation of the logarithmic gain of the training samples with non-silent masks. This is to teach the model to ignore the gain value of the silent mask. With the query, key, and value embeddings in place, any QKV-compatible attention block 1220a can be used to compute the output embeddings such that

.(11)(11)

然后,输出块1220b 计算预测分布,输出块1220b 预测分布的参数和。可以考虑三个类型的注意力,即加性注意力(additive attention,AA)、点积注意力(dot-product attention,DPA)和多头注意力(multi-head attention,MHA)。在该示例中,使用4头MHA。Then, output block 1220b computes the predicted distribution, and output block 1220b contains the parameters of the predicted distribution. Three types of attention can be considered: additive attention (AA), dot-product attention (DPA), and multi-head attention (MHA). In this example, 4-head MHA is used.

损失函数loss function

如同示例研究1,人类主观反应可以被认为是固有地随机的,一方面,参与者是总体的一个样本,另一方面,每个参与者在收听测试过程中给出的反应也是他们固有的非确定性感知的随机样本。因此,用于目标属性的标签(即ISO愉悦度)可以被视为代表目标属性分布的随机变量的观察结果,目标属性的真实分布是不明的。在使用最大似然计算式的情况下,PPAP模型的优化目标可以通过以下计算式给出:As in Example Study 1, subjective human responses can be considered inherently random. On the one hand, the participants are a sample of the population; on the other hand, each participant's response during the listening test is also a random sample of their inherently nondeterministic perceptions. Therefore, the label used for the target attribute (i.e., ISO pleasantness) can be viewed as an observation of a random variable representing the distribution of the target attribute, the true distribution of which is unknown. Using the maximum likelihood formula, the optimization objective of the PPAP model can be given by the following formula:

()(12)()(12)

其中是模型的参数,并且是在x处估算的X的对数密度。通过独立的正态分布对每个声学场景的主观响应进行建模,计算式(12)中的优化目标转换为损失函数Here, represents the model parameters, and represents the logarithmic density of X estimated at x. Subjective responses to each acoustic scene are modeled using independent normal distributions, and the optimization objective in equation (12) is transformed into a loss function.

,(13)(13)

其中是训练轮次。This includes training rounds.

优化推理Optimize reasoning

尽管模型在训练过程中看到不同声景-掩蔽器-增益样本的多个批次,但是在实时推理期间,模型每次只看到一个基本声景,与此同时,经历多个掩蔽器和增益级别以为特定声景找到最合适的掩蔽器-增益对。因此,通过消除重复计算,以代替为每个掩蔽器-增益组合进行新的前向传递,可以获得显著的性能改善。分别以、和表示执行、和所需的时间。设为每个声景待查询的掩蔽器的数量,并且为每个掩蔽器待查询的增益值的数量。朴素批处理方案将需要的总运行时间。相比之下,通过优化查询过程,如图13的算法1所示,总运行时间可以减少至。图13示出根据各个实施例的、减少总运行时间的算法。附加地,对于固定的掩蔽器组,可以预先计算掩蔽特征,这允许将总运行时间进一步减少至。Although the model sees multiple batches of different soundscape-mask-gain samples during training, during real-time inference, the model sees only one basic soundscape at a time, while simultaneously experiencing multiple maskers and gain levels to find the most suitable masker-gain pair for a specific soundscape. Therefore, a significant performance improvement can be achieved by eliminating redundant computations instead of performing a new forward pass for each masker-gain combination. Let , , and represent the execution time required for , , and , respectively. Let be the number of masks to be queried for each soundscape, and be the number of gain values to be queried for each mask. The naive batching scheme would require a total runtime of . In contrast, by optimizing the query process, as shown in Algorithm 1 of Figure 13, the total runtime can be reduced to . Figure 13 illustrates algorithms for reducing the total runtime according to various embodiments. Additionally, for a fixed set of maskers, masking features can be pre-computed, which allows for a further reduction in the total runtime to .

数据集Dataset

本文中使用的数据集从示例研究1中使用的数据集扩展有额外的参与者,共计442名参与者的18564个数据点(每位参与者42个刺激)。每个数据点代表参与者对增强声景的主观反应。The dataset used in this paper is an extension of the dataset used in Example Study 1, with additional participants, totaling 18,564 data points (42 stimuli per participant) for 442 participants. Each data point represents a participant's subjective response to the augmented soundscape.

A.刺激A. Stimulation

通过将30秒的“掩蔽器”添加至源自世界城市声景(USotW)数据库的声景的30秒立体声景录音,增强声景在5个互斥的折中制成。也包括没被增强的声景(即使用静音掩蔽器“增强”的声景)以作为对照。所使用的全部录音以44.1kHz采样。更多详细信息可以在示例研究1中找到。The enhanced soundscape was created by adding a 30-second “mask” to a 30-second stereo soundscape recording sourced from the City Soundscapes of the World (USotW) database, resulting in five mutually exclusive trade-offs. An unenhanced soundscape (i.e., one “enhanced” using a silence mask) was also included as a control. All recordings used were sampled at 44.1 kHz. More details can be found in Example Study 1.

将掩蔽器添加至声景是在{-6,-3,0,3,6}dBA的、特定的声景与掩蔽器的比率(SMR)下制成的,其中SMR由声景的SPL与掩蔽器的SPL的比率定义。虽然已知基本声景音轨的实时SPL,但是掩蔽器数据源自在线存储库,特别是Freesound和Xeno-canto,因此掩蔽器数据需要校准。通过在虚拟头(dummy head)(GRAS 45BB KEMAR Head & Torso)上以1dBA步长,对音轨进行彻底校准至46dBA至83dBA的闭区间中的SPL值,以创建将第j个掩蔽器放大至特定SPL所需的数字增益的查对表。校准过程是半自动的,其中使用与参与者相同的音频播放设备。由于基本声景的实时SPL通常是非整数的,非整数的数字增益是从查对表插值得到的,其中使用Masks were added to the soundscape at a specific soundscape-to-mask ratio (SMR) of {-6,-3,0,3,6} dBA, where the SMR is defined by the ratio of the soundscape's SPL to the mask's SPL. While the real-time SPL of the base soundscape track is known, the mask data is sourced from online repositories, specifically Freesound and Xeno-canto, and therefore requires calibration. A lookup table was created to amplify the j-th mask to a specific SPL by thoroughly calibrating the track to SPL values in a closed interval of 46 dBA to 83 dBA on a dummy head (GRAS 45BB KEMAR Head & Torso) in 1 dBA steps. The calibration process was semi-automatic, using the same audio playback devices as the participants. Since the real-time SPL of the base soundscape is typically non-integer, the non-integer digital gain was interpolated from the lookup table, where...

(14)(14)

其中是在dBA的SPL下用于掩蔽器的校准增益。This refers to the calibration gain used for the mask under the SPL of dBA.

B.主观反应B. Subjective reaction

全部参与者都使用一副环绕式耳机(Beyerdynamic Custom One Pro)收听经校准的增强声景,该耳机由外部声卡(Creative SoundBlaster E5)驱动。收听完每个增强声景后,参与者回答一份基于ISO12913-2标准的6项问卷:“您在多大程度上同意或反对当前周围声音环境是[...]?”,其中[...]是愉悦的、混乱的、活跃的、平静的、烦扰的和单调的中的一个。每个问卷项目都有“非常反对”、“反对”、“既不同意也不反对”、“同意”和“非常同意”的多个选项,该多个选项被数字编码为从1至5。然后通过以下计算式,计算每个增强声景的ISO愉悦度的基准真实值标签All participants listened to calibrated enhanced soundscapes using a pair of Beyerdynamic Custom One Pro surround sound headphones driven by an external sound card (Creative SoundBlaster E5). After listening to each enhanced soundscape, participants answered a six-item questionnaire based on the ISO 12913-2 standard: “To what extent do you agree or disagree with your current surrounding sound environment [...]?”, where [...] is one of pleasant, chaotic, active, calm, disturbing, or monotonous. Each questionnaire item had multiple options: “Strongly Disagree,” “Disagree,” “Neither Agree nor Disagree,” “Agree,” and “Strongly Agree,” which were numerically coded from 1 to 5. The baseline true value label of ISO pleasantness for each enhanced soundscape was then calculated using the following formula.

(15)(15)

其中、、、、、分别是愉悦的、烦扰的、平静的、混乱的、活跃的和单调的的问卷项目的数字评分。The numbers 1, 2, 3, 4, 5, 6, and 7 represent the numerical ratings of the questionnaire items, respectively, indicating pleasant, disturbing, calm, chaotic, active, and monotonous feelings.

测试test

对于本文中的全部模型类型,每个模型都以5折交叉验证的方式,为每个验证集的相同的10个种子,进行训练,每个模型类型总共50个模型。使用具有5×10-5的学习率的亚当优化器,对每个模型进行最多100轮的训练。如果验证均方误差(MSE)在至少5轮内没有下降,则学习率减半,并且如果验证MSE在至少10轮内没有下降,则提前停止训练。本文披露的MSE以及平均绝对误差(MAE)是使用预测分布的平均值与基准真实值标签之间的差值计算的。For all model types in this paper, each model was trained using 5-fold cross-validation with the same 10 seeds for each validation set, for a total of 50 models per model type. Each model was trained for a maximum of 100 epochs using an Adam optimizer with a learning rate of 5 × 10⁻⁵ . The learning rate was halved if the validation mean squared error (MSE) did not decrease within at least 5 epochs, and training was stopped early if the validation MSE did not decrease within at least 10 epochs. The MSE and mean absolute error (MAE) disclosed in this paper are calculated using the difference between the mean of the predicted distribution and the baseline true label.

除上文提到的三个注意力类型(在标题“模型”下)之外,测试中还包括作为基线模型的“直通”注意力块。基线模型在图14中以“X”表示。直通块可以定义为In addition to the three attention types mentioned above (under the heading "Model"), the test also included a "pass-through" attention block as the baseline model. The baseline model is represented by an "X" in Figure 14. A pass-through block can be defined as...

(16)(16)

图14示出(上)均方误差(MSE)作为注意力块类型的函数的图示,该图示出根据各个实施例的每个注意力块类型和特征增强方法的愉悦度预测的验证MSE的小提琴图;以及(下)平均绝对误差(MAE)作为注意力块类型的函数的图示,该图示出根据各个实施例的每个注意力块类型和特征增强方法的愉悦度预测的验证MAE的小提琴图。图14示出根据每个注意力和增强类型的MSE和MAE的预测误差分布。在全部注意力类型中,CONV增强方法在MSE和MAE两个方面表现最佳,其次是CAT和ADD。这可以归因于CONV方法和内核过滤,CONV方法天然地激励与之间的特征对齐(feature alignment),与其他两个方法相比,内核过滤允许在特征空间中进行较灵活的增强操作。Figure 14 shows (top) a graph of mean squared error (MSE) as a function of attention block type, illustrating a violin plot of the validation MSE for pleasure predictions for each attention block type and feature enhancement method according to various embodiments; and (bottom) a graph of mean absolute error (MAE) as a function of attention block type, illustrating a violin plot of the validation MAE for pleasure predictions for each attention block type and feature enhancement method according to various embodiments. Figure 14 shows the prediction error distribution of MSE and MAE for each attention and enhancement type. Among all attention types, the CONV enhancement method performs best in both MSE and MAE, followed by CAT and ADD. This can be attributed to the CONV method and kernel filtering. The CONV method naturally encourages feature alignment between attention and feature blocks, while kernel filtering allows for more flexible enhancement operations in the feature space compared to the other two methods.

还可以看出,与无注意力的模型相比,使用QKV注意力总体上允许获得较好的性能。然而,只要使用注意力的某个或某种形式将查询、键和值张量耦接,注意力机制的类型似乎不对模型性能产生显著影响。尽管AA和DPA只有一个注意力头,但是四头MHA似乎没有在预测误差方面获得可察觉的改善。还应当注意,特征增强块的选择似乎对模型的准确性有较显著的影响,如所看到的无注意力的CONV增强模型总体上比有注意力的ADD模型表现好的事实。It can also be seen that using QKV attention generally allows for better performance compared to models without attention. However, the type of attention mechanism does not seem to have a significant impact on model performance as long as some form of attention is used to couple the query, key, and value tensors. Although AA and DPA only have one attention head, the four-head MHA does not seem to achieve a noticeable improvement in prediction error. It should also be noted that the choice of feature augmentation blocks seems to have a significant impact on model accuracy, as evidenced by the fact that the non-attentional CONV augmentation model generally performs better than the attentional ADD model.

尽管模型从没看到呈现给参与者的、增强声景的音频数据,在0.122±0.005处的最佳平均MSE模型(具有CONV增强的卢昂注意力(Luong attention))与示例研究1中的结果相当,示例研究1使用增强声景的对数梅尔频谱图作为输入。这可以证明特征增强技术在代替波形空间的特征空间中模拟声景增强的有效性。Although the model never saw the audio data of the enhanced soundscape presented to the participants, the best-mean ...

图15A是ISO愉悦度作为对数增益的函数的图示,该图示出根据各个实施例,使用具有CONV增强(种子5、验证折2)的加性注意力(AA),为掩蔽器类别“水”进行增益插值。图15B是ISO愉悦度作为对数增益的函数的图示,该图示出根据各个实施例,使用具有CONV增强(种子5、验证折2)的加性注意力(AA),为掩蔽器类别“交通”进行增益插值。图15C是ISO愉悦度作为对数增益的函数的图示,该图示出根据各个实施例,使用具有CONV增强(种子5、验证折2)的加性注意力(AA),为掩蔽器类别“鸟”进行增益插值。图15D是ISO愉悦度作为对数增益的函数的图示,该图示出根据各个实施例,使用具有CONV增强(种子5、验证折2)的加性注意力(AA),为掩蔽器类别“施工(construction)”进行增益插值。图15E是ISO愉悦度作为对数增益的函数的图示,该图示出根据各个实施例,使用具有CONV增强(种子5、验证折2)的加性注意力(AA),为掩蔽器类别“静音”进行增益插值。为实现该可视化的目的,全部掩蔽器均在零对数增益下调整为65.0dBA。基本声景具有65.32dBA的实时SPL级别。在训练过程中,模型看不到基本声景和全部掩蔽器。图15A至图15E中的实线表示预测分布的平均值。阴影区表示最高至预测平均值上方和最低至预测平均值下方的预测标准差的愉悦度级别。Figure 15A is a diagram illustrating ISO pleasure as a function of logarithmic gain, showing gain interpolation for the masker category "Water" using additive attention (AA) with CONV enhancement (seed 5, verification fold 2) according to various embodiments. Figure 15B is a diagram illustrating ISO pleasure as a function of logarithmic gain, showing gain interpolation for the masker category "Traffic" using additive attention (AA) with CONV enhancement (seed 5, verification fold 2) according to various embodiments. Figure 15C is a diagram illustrating ISO pleasure as a function of logarithmic gain, showing gain interpolation for the masker category "Bird" using additive attention (AA) with CONV enhancement (seed 5, verification fold 2) according to various embodiments. Figure 15D is a diagram illustrating ISO pleasure as a function of logarithmic gain, showing gain interpolation for the masker category "Construction" using additive attention (AA) with CONV enhancement (seed 5, verification fold 2) according to various embodiments. Figure 15E is a graph illustrating ISO pleasantness as a function of logarithmic gain, showing gain interpolation for the masker category "Mute" using additive attention (AA) with CONV enhancements (seed 5, validation fold 2) according to various embodiments. For this visualization, all masks are tuned to 65.0 dBA at zero logarithmic gain. The base soundscape has a real-time SPL level of 65.32 dBA. During training, the model does not see the base soundscape or all masks. The solid lines in Figures 15A through 15E represent the mean of the prediction distribution. The shaded areas represent the pleasantness level from the highest to above the prediction mean and the lowest to below the prediction mean, representing the standard deviation of the prediction.

图15A至图15E可以通过为不同类型的掩蔽器在[-2,2]区间内、在256个对数增益值中查询ISO愉悦度分布,以展示模型的增益意识特性(gain-aware nature)。还包括静音“掩蔽器”音轨,以测试模型同时考虑掩蔽器和掩蔽器的增益级别二者的影响的能力。尽管愉悦度预测不是完全恒定的,可以看出,在不考虑查询增益值的情况下,模型都可以辨别出掩蔽器是静音的,并且输出大致恒定的愉悦度预测。对于其他非静音掩蔽器,可以看出,模型意识到每个掩蔽器的特性。与先前研究的发现一致,模型预测愉悦度随着鸟鸣掩蔽器增益的增加而增加,直到鸟鸣掩蔽器的增益接近环境级别(γ≈-0.5),然后愉悦度下降至没被增强的声景(即静音掩蔽器预测)的愉悦度以下。水声掩蔽器也呈现相似的效果,水声掩蔽器的、最高至γ≈−0.5的愉悦度级别高于没被增强的声景的愉悦度,但是由于水声较连续的特性,愉悦度级别在比鸟鸣掩蔽器低的增益处开始下降。对于已知令人不愉快的掩蔽器,例如交通和施工,随着增益级别的提高,模型正确地输出越来越令人不愉快的预测。Figures 15A through 15E demonstrate the model's gain-aware nature by querying the ISO pleasantness distribution across 256 logarithmic gain values in the [-2,2] interval for different types of maskers. A silent "masker" track is also included to test the model's ability to consider both the masker itself and its gain level. While the pleasantness predictions are not perfectly constant, it can be seen that the model can discern a silent masker without considering the query gain value and outputs a roughly constant pleasantness prediction. For other non-silent maskers, it can be seen that the model is aware of the characteristics of each masker. Consistent with previous findings, the model predicts pleasantness as the bird song masker gain increases until it approaches the ambient level (γ≈-0.5), after which the pleasantness drops below the pleasantness of the unenhanced soundscape (i.e., the silent masker prediction). Underwater acoustic maskers exhibit a similar effect, with the highest pleasantness level (up to γ≈−0.5) exceeding that of the unenhanced soundscape. However, due to the more continuous nature of underwater sounds, the pleasantness level begins to decrease at lower gains than with birdsong maskers. For maskers known to be unpleasant, such as traffic and construction, the model correctly outputs increasingly unpleasant predictions as the gain level increases.

结论in conclusion

各个实施例可以涉及改进的概率感知属性预测器模型,该模型允许对增强声景的主观响应进行增益意识预测。所提出的模型将掩蔽器与声景特征提取解耦,并且在特征空间中模拟声景增强过程,从而消除在波形域中进行计算方面成本高昂的混合过程的需求。此外,该模型被重制以考虑掩蔽器的数字增益级别,以代替声景与掩蔽器的比率,这允许在无需耗时的校准的情况下,进行源自多个源的掩蔽器的使用。模型的模块化设计允许在推理时间内,进行大量特征的重用和预计算,这减少了部署所需的总体延迟和计算资源。在使用源自442名参与者的18000个主观响应的大规模数据集的情况下,已经证明了该模型能够准确预测与声景、掩蔽器和增益相关的愉悦度评分的能力。Various embodiments may involve an improved probabilistically aware attribute predictor model that allows for gain-aware prediction of subjective responses to enhanced soundscapes. The proposed model decouples the masker from soundscape feature extraction and models the soundscape enhancement process in the feature space, thereby eliminating the need for computationally costly mixing processes in the waveform domain. Furthermore, the model is refactored to consider the digital gain level of the masker instead of the soundscape-to-mask ratio, allowing for the use of masks from multiple sources without time-consuming calibration. The modular design of the model allows for the reuse and pre-computation of a large number of features during inference time, reducing the overall latency and computational resources required for deployment. The model has been demonstrated to accurately predict pleasure ratings associated with soundscape, masker, and gain using a large dataset of 18,000 subjective responses from 442 participants.

在声景增强系统中,掩蔽器和播放增益级别的选择对于声景增强系统在改善给定环境的整体声学舒适度的有效性而言是至关重要的。常规上,适当的掩蔽器和增益级别的选择是由可能不能代表目标人群的专家意见决定的,或者是由可能耗时耗力的收听测试决定的。此外,所获得的掩蔽器和增益的静态选择通常对于现实世界声景的动态特性而言是不灵活的。在本文中,深度学习模型被用于对给定声景的最佳掩蔽器及最佳掩蔽器的增益级别进行联合选择。采用高度模块化的构建块设计所提出的模型,这允许获得优化的推理过程,优化的推理过程可以快速搜索大量掩蔽器和增益的组合。此外,引入在数字增益级别方面调节的特征域声景增强的使用,这消除了推理期间计算方面的成本高昂的波形域混合过程,以及掩蔽器所需的、繁琐的预校准过程。所提出的系统在超过440名参与者对增强声景的主观响应的大规模数据集上得到验证,这确保了模型能够预测掩蔽器及掩蔽器的增益级别对感知愉悦度级别的综合影响的能力。In soundscape enhancement systems, the selection of maskers and playback gain levels is crucial to the effectiveness of the system in improving the overall acoustic comfort of a given environment. Conventionally, the selection of appropriate maskers and gain levels is determined by expert opinions that may not be representative of the target audience, or by potentially time-consuming and laborious listening tests. Furthermore, the resulting static selections of maskers and gains are often inflexible in relation to the dynamic characteristics of real-world soundscapes. In this paper, a deep learning model is used to jointly select the optimal masker and the optimal gain level for a given soundscape. The proposed model employs a highly modular building block design, allowing for an optimized inference process that can rapidly search a large number of masker and gain combinations. Moreover, the introduction of feature-domain soundscape enhancement with tuning at the digital gain level eliminates the computationally costly waveform-domain mixing process during inference, as well as the cumbersome pre-calibration process required for maskers. The proposed system was validated on a large dataset of subjective responses to enhanced soundscapes from more than 440 participants, which ensures the model’s ability to predict the combined effect of the mask and the mask’s gain level on the perceived pleasure level.

Claims (20)

1.声景增强系统,包括:1. Soundscape enhancement system, including: 数据获取系统,所述数据获取系统被设置为提供周围声景数据;A data acquisition system is configured to provide ambient sound data. 数据库,所述数据库包括多个掩蔽设置;The database includes multiple masking settings; 感知属性预测器,所述感知属性预测器耦接于所述数据获取系统和所述数据库,所述感知属性预测器被设置为基于所述周围声景数据,为所述多个掩蔽设置中的每个掩蔽设置,生成在一个或多个预定义感知属性指标上代表感知的预测;A perception attribute predictor, coupled to the data acquisition system and the database, is configured to generate a prediction representing perception on one or more predefined perception attribute metrics for each of the plurality of masking settings, based on the surrounding soundscape data. 掩蔽设置排序系统,所述掩蔽设置排序系统被设置为基于所述感知属性预测器生成的所述预测,确定一个或多个最佳掩蔽设置;以及A masking setting ranking system, configured to determine one or more optimal masking settings based on the predictions generated by the perceptual attribute predictor; and 播放系统,所述播放系统被设置为播放或再现所述一个或多个最佳掩蔽设置。A playback system configured to play or reproduce the one or more optimal masking settings. 2.根据权利要求1所述的声景增强系统,其中所述数据获取系统还被设置为提供群体统计或情境数据。2. The soundscape enhancement system according to claim 1, wherein the data acquisition system is further configured to provide crowd statistics or contextual data. 3.根据权利要求2所述的声景增强系统,其中所述感知属性预测器被设置为还基于所述群体统计或情境数据,生成所述预测。3. The soundscape enhancement system of claim 2, wherein the perceptual attribute predictor is configured to generate the prediction based on the group statistics or contextual data. 4.根据权利要求2至3中任一项所述的声景增强系统,其中所述群体统计或情境数据包括与环境参数相关的数据、一名或多名听众的群体统计数据、所述一名或多名听众的心理数据、所述一名或多名听众的生理数据或其任意组合。4. The soundscape enhancement system according to any one of claims 2 to 3, wherein the group statistics or contextual data includes data related to environmental parameters, group statistics of one or more listeners, psychological data of one or more listeners, physiological data of one or more listeners, or any combination thereof. 5.根据权利要求4所述的声景增强系统,其中所述与环境参数相关的数据包括与位置相关的数据、与所述视觉环境相关的数据、附近的人数、空气温度、湿度、风速或其任意组合。5. The soundscape enhancement system according to claim 4, wherein the environmental parameter-related data includes location-related data, visual environment-related data, the number of people nearby, air temperature, humidity, wind speed, or any combination thereof. 6.根据权利要求4或5所述的声景增强系统,其中所述一名或多名听众的所述群体统计数据包括与年龄、性别、职业或其任意组合相关的数据。6. The soundscape enhancement system according to claim 4 or 5, wherein the group statistics of the one or more listeners include data related to age, gender, occupation, or any combination thereof. 7.根据权利要求4至6中任一项所述的声景增强系统,其中所述一名或多名听众的所述心理数据包括与噪声敏感度、感知压力、幸福感指数评分或其任意组合相关的数据。7. The soundscape enhancement system according to any one of claims 4 to 6, wherein the psychological data of the one or more listeners includes data related to noise sensitivity, perceived stress, well-being index scores, or any combination thereof. 8.根据权利要求4至7中任一项所述的声景增强系统,其中所述一名或多名听众的所述生理数据包括与心率、血压、体温或其任意组合相关的数据。8. The soundscape enhancement system according to any one of claims 4 to 7, wherein the physiological data of the one or more listeners includes data related to heart rate, blood pressure, body temperature or any combination thereof. 9.根据权利要求1至8中任一项所述的声景增强系统,其中所述感知属性预测器被设置为基于一个或多个掩蔽增益输入,生成预测。9. The soundscape enhancement system according to any one of claims 1 to 8, wherein the perceptual attribute predictor is configured to generate predictions based on one or more masking gain inputs. 10.根据权利要求1至9中任一项所述的声景增强系统,其中所述一个或多个预定义感知属性指标包括愉悦度、活跃度、事件性、平静度、感知响度、声音质量、锐度、粗糙度或其任意组合。10. The soundscape enhancement system according to any one of claims 1 to 9, wherein the one or more predefined perceptual attribute indicators include pleasantness, activity, eventfulness, calmness, perceived loudness, sound quality, sharpness, roughness, or any combination thereof. 11.根据权利要求1至10中任一项所述的声景增强系统,其中所述感知属性预测器被设置为通过组合所述周围声景数据和每个掩蔽设置,以生成所述预测。11. The soundscape enhancement system according to any one of claims 1 to 10, wherein the perceptual attribute predictor is configured to generate the prediction by combining the surrounding soundscape data and each masking setting. 12.根据权利要求1至11中任一项所述的声景增强系统,12. The soundscape enhancement system according to any one of claims 1 to 11, 其中所述感知属性预测器包括设置为提取所述周围声景数据的特征的声景特征提取器;The perceptual attribute predictor includes a sound scene feature extractor configured to extract features from the surrounding sound scene data. 其中所述感知属性预测器包括设置为提取每个掩蔽设置的特征的掩蔽特征提取器;并且The perceptual attribute predictor includes a masking feature extractor configured to extract features for each masking setting; and 其中所述感知属性预测器被设置为基于所述周围声景数据的提取特征和每个掩蔽设置的提取特征,生成所述预测。The perceptual attribute predictor is configured to generate the prediction based on the extracted features of the surrounding soundscape data and the extracted features of each masking setting. 13.根据权利要求1至12中任一项所述的声景增强系统,13. The soundscape enhancement system according to any one of claims 1 to 12, 其中所述感知属性预测器包括设置为从所述周围声景数据和每个掩蔽设置提取客观特征的客观特征块;The perceptual attribute predictor includes objective feature blocks configured to extract objective features from the surrounding acoustic scene data and each masking setting; 其中所述感知属性预测器包括设置为从所述群体统计或情境数据提取主观特征的主观特征块;并且The perceptual attribute predictor includes a subjective feature block configured to extract subjective features from the group statistics or contextual data; and 其中所述感知属性预测器被设置为基于所述提取的客观特征和所述提取的主观特征,生成所述预测。The perceptual attribute predictor is configured to generate the prediction based on the extracted objective features and the extracted subjective features. 14.根据权利要求1至13中任一项所述的声景增强系统,其中所述感知属性预测器包括线性回归模型,所述模型线性回归模型被设置为基于声学或心理声学参数,生成所述预测,所述声学或心理声学参数是基于所述周围声景数据和每个掩蔽设置计算的。14. The soundscape enhancement system according to any one of claims 1 to 13, wherein the perceptual attribute predictor comprises a linear regression model, the linear regression model being configured to generate the prediction based on acoustic or psychoacoustic parameters calculated based on the surrounding soundscape data and each masking setting. 15.根据权利要求1至14中任一项所述的声景增强系统,其中所述感知属性预测器包括一个或多个深度神经网络,所述一个或多个深度神经网络被设置为基于原始音频、基于所述周围声景数据和每个掩蔽设置计算的频谱图呈现,或者所述原始音频和所述频谱图呈现的组合,生成所述预测。15. The soundscape enhancement system according to any one of claims 1 to 14, wherein the perceptual attribute predictor comprises one or more deep neural networks configured to generate the prediction based on raw audio, spectrogram rendering calculated based on the surrounding soundscape data and each masking setting, or a combination of the raw audio and the spectrogram rendering. 16.根据权利要求1至15中任一项所述的声景增强系统,其中所述数据获取系统被设置为基于一个或多个收音器、一根或多根天线、存储介质或其任意组合的输入,生成所述周围声景数据。16. The soundscape enhancement system according to any one of claims 1 to 15, wherein the data acquisition system is configured to generate the surrounding soundscape data based on inputs from one or more microphones, one or more antennas, a storage medium or any combination thereof. 17.根据权利要求1至16中任一项所述的声景增强系统,其中所述多个掩蔽设置包括一个或多个录制的或合成的声音的音轨、一个或多个静音的音轨,或者从所述一个或多个录制的或合成的声音的音轨和所述一个或多个静音的音轨衍生出的一个或多个音轨。17. The soundscape enhancement system according to any one of claims 1 to 16, wherein the plurality of masking settings comprises one or more recorded or synthesized sound tracks, one or more silent sound tracks, or one or more sound tracks derived from the one or more recorded or synthesized sound tracks and the one or more silent sound tracks. 18.根据权利要求1至17中任一项所述的声景增强系统,其中所述播放系统包括一个或多个扬声器、一个或多个虚拟现实或增强现实耳麦、一个或多个耳机、一个或多个耳塞或其任意组合。18. The soundscape enhancement system according to any one of claims 1 to 17, wherein the playback system comprises one or more speakers, one or more virtual reality or augmented reality headsets, one or more headphones, one or more earbuds, or any combination thereof. 19.根据权利要求1至18中任一项所述的声景增强系统,其中生成的所述预测是不确定性的。19. The soundscape enhancement system according to any one of claims 1 to 18, wherein the generated prediction is uncertain. 20.一种声景增强系统的形成方法,所述方法包括:20. A method for forming a soundscape enhancement system, the method comprising: 提供数据获取系统,所述数据获取系统设置为提供周围声景数据;A data acquisition system is provided, wherein the data acquisition system is configured to provide ambient sound scene data; 提供数据库,所述数据库包括多个掩蔽设置;Provide a database that includes multiple masking settings; 将感知属性预测器耦接于所述数据获取系统和所述数据库,所述感知属性预测器被设置为基于所述周围声景数据,为所述多个掩蔽设置中的每个掩蔽设置生成在一个或多个预定义感知属性指标上代表感知的预测;A perception attribute predictor is coupled to the data acquisition system and the database. The perception attribute predictor is configured to generate a prediction representing perception on one or more predefined perception attribute metrics for each of the plurality of masking settings based on the surrounding soundscape data. 提供掩蔽设置排序系统,所述掩蔽设置排序系统被设置为基于所述感知属性预测器生成的所述预测,确定一个或多个最佳掩蔽设置;以及Provide a masking setup ranking system, the masking setup ranking system being configured to determine one or more optimal masking setups based on the predictions generated by the perceptual attribute predictor; and 提供播放系统,所述播放系统设置为播放或再现所述一个或多个最佳掩蔽设置。A playback system is provided, which is configured to play or reproduce the one or more optimal masking settings.
HK62024100927.8A 2022-04-27 2023-04-26 Soundscape augmentation system and method of forming the same HK40112923A (en)

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
SG10202204451S 2022-04-27

Publications (1)

Publication Number Publication Date
HK40112923A true HK40112923A (en) 2025-01-28

Family

ID=

Similar Documents

Publication Publication Date Title
Best et al. Sound externalization: A review of recent research
Neidhardt et al. Perceptual matching of room acoustics for auditory augmented reality in small rooms-literature review and theoretical framework
Brambilla et al. Merging physical parameters and laboratory subjective ratings for the soundscape assessment of urban squares
US9319019B2 (en) Method for augmenting a listening experience
US20160234606A1 (en) Method for augmenting hearing
CN109076305A (en) The rendering of augmented reality earphone environment
CN104052423A (en) Custom Audio Reproduction Devices
Ntalampiras A transfer learning framework for predicting the emotional content of generalized sound events
WO2023211385A1 (en) Soundscape augmentation system and method of forming the same
Lundén et al. On urban soundscape mapping: A computer can predict the outcome of soundscape assessments
CN120418861A (en) Generating digital media based on blockchain data
US20190261102A1 (en) Remotely updating a hearing aid profile
Lopez-Ballester et al. AI-IoT platform for blind estimation of room acoustic parameters based on deep neural networks
Watcharasupat et al. Autonomous in-situ soundscape augmentation via joint selection of masker and gain
WO2023179765A1 (en) Multimedia recommendation method and apparatus
US11432078B1 (en) Method and system for customized amplification of auditory signals providing enhanced karaoke experience for hearing-deficient users
Brambilla et al. Measurements and techniques in soundscape research
HK40112923A (en) Soundscape augmentation system and method of forming the same
JP3657261B2 (en) Hearing test system and hearing aid selection system using the same
JP3106663U (en) Hearing test system and hearing aid selection system using the same
Grimm et al. Virtual acoustic environments for comprehensive evaluation of model-based hearing devices
US11490218B1 (en) Time domain neural networks for spatial audio reproduction
WO2022229287A1 (en) Methods and devices for hearing training
Ortiz Auditorio 400 at the ‘Museo Reina Sofia’in Madrid: Use of variable systems for acoustic improvements
US12587798B2 (en) Headphones with sound-enhancement and integrated self-administered hearing test