CN115291718A - Man-machine interaction system in smart home space and application method thereof - Google Patents
Man-machine interaction system in smart home space and application method thereof Download PDFInfo
- Publication number
- CN115291718A CN115291718A CN202210861645.9A CN202210861645A CN115291718A CN 115291718 A CN115291718 A CN 115291718A CN 202210861645 A CN202210861645 A CN 202210861645A CN 115291718 A CN115291718 A CN 115291718A
- Authority
- CN
- China
- Prior art keywords
- information
- intelligent
- perception
- sensor
- user
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/70—Multimodal biometrics, e.g. combining information from different biometric modalities
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F2203/00—Indexing scheme relating to G06F3/00 - G06F3/048
- G06F2203/01—Indexing scheme relating to G06F3/01
- G06F2203/011—Emotion or mood input determined on the basis of sensed human body parameters such as pulse, heart rate or beat, temperature of skin, facial expressions, iris, voice pitch, brain activity patterns
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Human Computer Interaction (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- Multimedia (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
本发明涉及一种智能家居空间下的人机交互系统及其应用方法,系统包括部署于智能家居空间中的多通道传感器网络、智能感知模块和人机交互模块;多通道传感器网络用于对智能家居空间进行多模态信息采集;智能感知模块用于根据采集的多模态信息对用户进行智能感知;人机交互模块用于基于用户智能感知结果实现各种智能终端设备与用户的多模态人机交互;其中,智能感知至少包括行为感知和情感感知。本发明利用多通道传感器网络采集智能家居环境下的多源异构数据,并利用多模态融合分析方式从多源异构数据中精准感知用户自然行为和情感状态;在此基础上实现智能家居空间中多模态人机交互,进而为用户提供更加智能的家居环境。
The invention relates to a human-computer interaction system in a smart home space and an application method thereof. The system includes a multi-channel sensor network, an intelligent perception module and a human-computer interaction module deployed in the smart home space; the multi-channel sensor network is used for intelligent The home space is used for multi-modal information collection; the intelligent perception module is used to intelligently perceive the user according to the collected multi-modal information; the human-computer interaction module is used to realize the multi-modal communication between various intelligent terminal devices and users based on the user's intelligent perception results Human-computer interaction; among them, intelligent perception includes at least behavior perception and emotion perception. The invention uses multi-channel sensor network to collect multi-source heterogeneous data in the smart home environment, and uses multi-modal fusion analysis method to accurately perceive the user's natural behavior and emotional state from the multi-source heterogeneous data; on this basis, the smart home is realized Multi-modal human-computer interaction in space, thereby providing users with a more intelligent home environment.
Description
技术领域technical field
本发明涉及智能家居技术领域,尤其涉及一种智能家居空间下的人机交互系统及其应用方法。The present invention relates to the technical field of smart home, in particular to a human-computer interaction system in a smart home space and an application method thereof.
背景技术Background technique
智能家居是通过物联网技术将家中的各种家居设备连接到一起,提供家居设备的智能联网控制。与普通家居相比,智能家居不仅具有传统的居住功能,还具有家居设备信息共享和自动化等特点,可以为家庭创造高品质的生活环境。Smart home is to connect various household devices in the home through the Internet of Things technology to provide intelligent networking control of household devices. Compared with ordinary homes, smart homes not only have traditional living functions, but also have the characteristics of information sharing and automation of home equipment, which can create a high-quality living environment for families.
然而,现有智能家居,形态受限于其技术手段和产业结构,只能以独立家居智能产品的组合形式为主,缺乏体系化、全局化、统一操控的智能设备互联解决方案及其关键技术。这种智能家居仅可称为“单品智能”家居,而非“整体智能”家居,其在人机交互时缺乏对用户的深度理解,无法实现真正的智能感知。However, the form of the existing smart home is limited by its technical means and industrial structure. It can only be a combination of independent home smart products, lacking a systematic, global, and unified control of smart device interconnection solutions and key technologies. . This kind of smart home can only be called a "single-product smart" home, not an "overall smart" home. It lacks a deep understanding of users during human-computer interaction and cannot achieve true intelligent perception.
因此,本发明致力于打破现有的智能产品孤岛模式,由单体“智能”向整体“智慧”转型,构造一整套以用户为中心的智能化、个性化、人性化的智能家居产品。Therefore, the present invention is committed to breaking the existing island model of smart products, transforming from single "smart" to overall "smart", and constructing a complete set of user-centered intelligent, personalized and humanized smart home products.
发明内容Contents of the invention
本发明的目的是提供一种智能家居空间下的人机交互系统及其应用方法,解决现有智能家居在人机交互时缺乏对用户的深度理解,无法实现真正的智能感知的问题,为用户带来家居环境和家居实体相融合的智慧、舒适、温馨、便捷的个性化家居体验。The purpose of the present invention is to provide a human-computer interaction system and its application method in the smart home space, to solve the problem that the existing smart home lacks a deep understanding of the user during human-computer interaction, and cannot realize real intelligent perception, and provide users with Bring a smart, comfortable, warm and convenient personalized home experience that integrates home environment and home entities.
第一方面,本发明提供一种智能家居空间下的人机交互系统,所述系统包括:In a first aspect, the present invention provides a human-computer interaction system in a smart home space, the system comprising:
所述多通道传感器网络,用于对所述智能家居空间进行多模态信息采集;The multi-channel sensor network is used to collect multi-modal information on the smart home space;
所述智能感知模块,用于根据采集的多模态信息对所述智能家居空间中的用户进行智能感知;The intelligent perception module is configured to perform intelligent perception on users in the smart home space according to the collected multimodal information;
所述人机交互模块,用于基于用户智能感知结果,实现所述智能家居空间中的各种智能终端设备与所述用户之间的多模态人机交互;The human-computer interaction module is configured to realize multimodal human-computer interaction between various intelligent terminal devices in the smart home space and the user based on the user's intelligent perception result;
其中,所述智能感知至少包括行为感知和情感感知。Wherein, the intelligent perception includes at least behavior perception and emotion perception.
根据本发明提供的智能家居空间下的人机交互系统,所述多模态信息包括图像信息、语音信息、触觉感知信息、气味信息和环境信息;所述多通道传感器网络包括:视觉通道、听觉通道、触觉通道、嗅觉通道和环境信息通道;According to the human-computer interaction system under the smart home space provided by the present invention, the multimodal information includes image information, voice information, tactile perception information, smell information and environmental information; the multi-channel sensor network includes: visual channel, auditory channels, tactile channels, olfactory channels and environmental information channels;
所述视觉通道包括视觉传感器,用于利用所述视觉传感器采集所述图像信息;The visual channel includes a visual sensor for collecting the image information using the visual sensor;
所述听觉通道包括听觉传感器,用于利用所述听觉传感器采集所述语音信息;The auditory channel includes an auditory sensor for collecting the voice information by using the auditory sensor;
所述触觉通道包括触觉传感器,用于利用所述触觉传感器采集所述触觉感知信息;The tactile channel includes a tactile sensor for collecting the tactile perception information by using the tactile sensor;
所述嗅觉通道包括嗅觉传感器,用于利用所述嗅觉传感器采集所述气味信息;The olfactory channel includes an olfactory sensor for collecting the odor information by using the olfactory sensor;
所述环境信息通道包括环境探测传感器,用于利用所述环境探测传感器采集所述环境信息。The environment information channel includes an environment detection sensor for collecting the environment information by using the environment detection sensor.
根据本发明提供的智能家居空间下的人机交互系统,所述视觉传感器部署于智能家居空间的房顶,至少包括以下一种:According to the human-computer interaction system under the smart home space provided by the present invention, the visual sensor is deployed on the roof of the smart home space, and at least includes the following one:
RGB相机矩阵、360度全景相机、红外相机和深度相机;RGB camera matrix, 360-degree panoramic camera, infrared camera and depth camera;
所述触觉传感器至少包括以下一种:The tactile sensor includes at least one of the following:
电容地板传感器、布置于智能沙发和/或智能床内的压力传感器;Capacitive floor sensors, pressure sensors arranged in smart sofas and/or smart beds;
所述听觉传感器包括:麦克风阵列;The auditory sensor includes: a microphone array;
所述环境信息传感器至少包括以下一种:The environmental information sensor includes at least one of the following:
温度传感器、湿度传感器、布置于智能床内的二氧化碳传感器、空气质量传感器;Temperature sensor, humidity sensor, carbon dioxide sensor and air quality sensor arranged in the smart bed;
其中,所述电容地板传感器中每个地板面片的电容值由其承受的压力决定。Wherein, the capacitance value of each floor patch in the capacitive floor sensor is determined by the pressure it bears.
根据本发明提供的智能家居空间下的人机交互系统,所述多通道传感器网络,还包括:触发设备;According to the human-computer interaction system under the smart home space provided by the present invention, the multi-channel sensor network further includes: a trigger device;
所述触发设备,用于以一定的频率触发所述视觉通道、所述听觉通道、所述触觉通道、所述嗅觉通道和所述环境信息通道进行信息的同步采集。The triggering device is configured to trigger the visual channel, the auditory channel, the tactile channel, the olfactory channel and the environmental information channel at a certain frequency to perform synchronous collection of information.
根据本发明提供的智能家居空间下的人机交互系统,所述智能感知模块包括:预处理单元、多模态特征提取单元和多模态特征融合解析单元;According to the human-computer interaction system under the smart home space provided by the present invention, the intelligent perception module includes: a preprocessing unit, a multimodal feature extraction unit and a multimodal feature fusion analysis unit;
所述预处理单元,用于对所述图像信息、所述语音信息、所述触觉感知信息、所述气味信息和所述环境信息进行预处理;The preprocessing unit is configured to preprocess the image information, the voice information, the tactile perception information, the smell information and the environmental information;
所述多模态特征提取单元,用于对预处理后的图像信息、语音信息、触觉感知信息、气味信息和环境信息进行特征提取,得到多模态特征信息;The multimodal feature extraction unit is used to perform feature extraction on the preprocessed image information, voice information, tactile perception information, odor information and environmental information to obtain multimodal feature information;
所述多模态特征融合解析单元,用于对所述多模态特征信息进行融合分析,确定所述用户智能感知结果。The multimodal feature fusion analysis unit is configured to perform fusion analysis on the multimodal feature information to determine the user's intelligent perception result.
根据本发明提供的智能家居空间下的人机交互系统,对所述图像信息进行预处理,至少包括以下一种:According to the human-computer interaction system under the smart home space provided by the present invention, the preprocessing of the image information includes at least one of the following:
对所述图像信息进行数据对齐;对所述图像信息进行去冗余;对所述图像信息进行除噪;performing data alignment on the image information; performing de-redundancy on the image information; denoising the image information;
对所述语音信息进行预处理,至少包括以下一种:Preprocessing the voice information includes at least one of the following:
对所述语音信息进行背景音和人声分离;去除所述语音信息中的噪声;去除所述语音信息中的空白音;performing background sound and human voice separation on the voice information; removing noise in the voice information; removing blank tones in the voice information;
对所述触觉感知信息进行预处理,至少包括以下一种:Preprocessing the tactile perception information includes at least one of the following:
对所述触觉感知信息中明显错误数据进行修正;Correcting the obviously wrong data in the tactile perception information;
对所述触觉感知信息中缺失数据进行补全;Completing missing data in the tactile perception information;
对所述气味信息进行预处理,至少包括以下一种:Preprocessing the odor information includes at least one of the following:
去除所述气味信息中包含的辨识失败信息;removing the identification failure information contained in the odor information;
去除所述气味信息中与所述用户的行为习惯明显背离的信息;removing information that obviously deviates from the user's behavior habits in the smell information;
对所述环境信息进行预处理,至少包括以下一种:Preprocessing the environmental information includes at least one of the following:
对所述环境信息中的缺失数据进行补全;Complete missing data in the environmental information;
对所述环境信息中因采集频率所丢失的数据使用插值的方法进行补全。The interpolation method is used to supplement the data lost due to the collection frequency in the environmental information.
根据本发明提供的智能家居空间下的人机交互系统,所述多模态特征提取单元,包括:姿态特征提取子单元、人脸特征提取子单元、热成像特征提取子单元、深度特征提取子单元、语义特征提取子单元、位姿特征提取子单元、用户睡眠质量特征提取子单元、用户运动轨迹特征提取子单元、用户行为事件特征提取子单元、温度特征提取子单元、湿度特征提取子单元和空气质量特征提取子单元;According to the human-computer interaction system under the smart home space provided by the present invention, the multimodal feature extraction unit includes: a posture feature extraction subunit, a face feature extraction subunit, a thermal imaging feature extraction subunit, and a depth feature extraction subunit. Unit, Semantic Feature Extraction Subunit, Pose Feature Extraction Subunit, User Sleep Quality Feature Extraction Subunit, User Motion Trajectory Feature Extraction Subunit, User Behavior Event Feature Extraction Subunit, Temperature Feature Extraction Subunit, Humidity Feature Extraction Subunit and air quality feature extraction subunit;
其中,所述姿态特征提取子单元,用于将所述RGB相机矩阵、360度全景相机、处于黑暗环境下的红外相机和/或所述深度相机采集的图像信息经预处理后的数据作为第一目标数据,并基于姿态识别算法从所述第一目标数据中提取姿态特征;Wherein, the posture feature extraction subunit is used to use the preprocessed data of the RGB camera matrix, the 360-degree panoramic camera, the infrared camera in a dark environment and/or the image information collected by the depth camera as the first a target data, and extract gesture features from the first target data based on a gesture recognition algorithm;
所述人脸特征提取子单元,用于将所述RGB相机矩阵、360度全景相机采集的图像信息经预处理后的数据作为第二目标数据,并基于人脸识别算法从所述第二目标数据中提取人脸特征;The human face feature extraction subunit is used to use the preprocessed data of the image information collected by the RGB camera matrix and the 360-degree panoramic camera as the second target data, and extract from the second target data based on the face recognition algorithm. Extract facial features from the data;
所述热成像特征提取子单元,用于将所述红外相机采集的图像信息经预处理后的数据作为第三目标数据,并从所述第三目标数据中提取人体热成像特征以及物体热成像特征;The thermal imaging feature extraction subunit is used to use the preprocessed data of the image information collected by the infrared camera as the third target data, and extract the thermal imaging features of the human body and the thermal imaging of the object from the third target data feature;
所述深度特征提取子单元,用于将所述深度相机采集的图像信息经预处理后的数据作为第四目标数据,并利用深度学习算法或三维点云特征点提取算法从所述第四目标数据中提取人体深度特征、空间深度特征和物体深度特征;The depth feature extraction subunit is used to use the preprocessed data of the image information collected by the depth camera as the fourth target data, and use a deep learning algorithm or a three-dimensional point cloud feature point extraction algorithm to extract the data from the fourth target Extract human body depth features, spatial depth features and object depth features from the data;
所述语义特征提取子单元,用于将所述麦克风阵列采集的语音信息经预处理后的数据作为第五目标数据,并利用语音识别算法和/或语音情感分析算法,从所述第五目标数据中提取语义特征;The semantic feature extraction subunit is configured to use the preprocessed data of the voice information collected by the microphone array as the fifth target data, and use a voice recognition algorithm and/or a voice emotion analysis algorithm to extract the data from the fifth target Extract semantic features from the data;
所述位姿特征提取子单元,用于将所述布置于智能沙发和/或智能床内的压力传感器采集的压力值信息经预处理后的数据作为第六目标数据,并分析所述第六目标数据得到用户在所述智能沙发和/或智能床上的位姿特征;The pose feature extraction subunit is used to use the preprocessed data of the pressure value information collected by the pressure sensor arranged in the smart sofa and/or smart bed as the sixth target data, and analyze the sixth target data. The target data obtains the pose characteristics of the user on the smart sofa and/or smart bed;
所述用户运动轨迹特征提取子单元,用于将所述电容地板传感器采集的电容值信息经预处理后的数据作为第七目标数据,并分析所述第七目标数据得到用户运动轨迹特征;所述用户运动轨迹特征包括:位置、步态、方向和速度;The user movement trajectory feature extraction subunit is used to use the preprocessed data of the capacitance value information collected by the capacitive floor sensor as the seventh target data, and analyze the seventh target data to obtain the user movement trajectory characteristics; The characteristics of the user's motion trajectory include: position, gait, direction and speed;
所述用户行为事件特征提取子单元,用于将所述嗅觉传感器采集的气味信息经预处理后的数据作为第八目标数据,并分析所述第八目标数据得到用户行为事件特征;The user behavior event feature extraction subunit is used to use the preprocessed data of the smell information collected by the olfactory sensor as the eighth target data, and analyze the eighth target data to obtain the user behavior event features;
所述用户睡眠质量特征子提取单元,用于将所述布置于智能床内的二氧化碳传感器采集的二氧化碳值信息经预处理后的数据作为第九目标数据,并分析所述第九目标数据得到用户睡眠质量特征;The user sleep quality feature sub-extraction unit is used to use the preprocessed data of the carbon dioxide value information collected by the carbon dioxide sensor arranged in the smart bed as the ninth target data, and analyze the ninth target data to obtain the user sleep quality characteristics;
所述温度特征提取子单元,用于将所述温度传感器采集的温度信息经预处理后的数据作为第十目标数据,并分析所述第十目标数据得到温度特征;The temperature feature extraction subunit is configured to use the preprocessed data of temperature information collected by the temperature sensor as the tenth target data, and analyze the tenth target data to obtain a temperature feature;
所述湿度特征提取子单元,用于将所述湿度传感器采集的湿度信息经预处理后的数据作为第十一目标数据,并分析所述第十一目标数据得到湿度特征;The humidity feature extraction subunit is configured to use the preprocessed data of the humidity information collected by the humidity sensor as the eleventh target data, and analyze the eleventh target data to obtain a humidity feature;
所述空气质量特征提取子单元,用于将所述空气质量传感器采集的空气质量信息经预处理后的数据作为第十二目标数据,并分析所述第十二目标数据得到空气质量特征。The air quality feature extraction subunit is configured to use the preprocessed data of the air quality information collected by the air quality sensor as the twelfth target data, and analyze the twelfth target data to obtain the air quality feature.
根据本发明提供的智能家居空间下的人机交互系统,所述多模态特征融合解析单元,包括:数据存储单元、智能感知模型构造单元和智能感知执行单元;According to the human-computer interaction system under the smart home space provided by the present invention, the multimodal feature fusion analysis unit includes: a data storage unit, an intelligent perception model construction unit and an intelligent perception execution unit;
所述数据存储单元,用于将所述多模态特征中各个模态特征存储于各个模态特征数据库中;The data storage unit is used to store each modal feature in the multi-modal feature in each modal feature database;
所述智能感知模型构造单元,用于在训练阶段利用各个模态特征数据库训练基于各个模态特征的智能感知模型;并将所述基于各个模态特征的智能感知模型进行模型融合得到智能感知模型;还用于在应用阶段利用所述多模态特征更新所述智能感知模型;The intelligent perception model construction unit is used to train the intelligent perception model based on each modal feature using each modal feature database in the training phase; and the intelligent perception model based on each modal feature is model-fused to obtain the intelligent perception model ; Also used to update the intellisense model using the multi-modal feature in the application phase;
所述智能感知执行单元,用于将所述多模态特征输入所述智能感知模型,得到所述智能感知模型输出的用户智能感知结果。The intelligent perception executing unit is configured to input the multimodal features into the intelligent perception model, and obtain the user's intelligent perception result output by the intelligent perception model.
根据本发明提供的智能家居空间下的人机交互系统,所述模型融合方式,包括但不限于:Voting、Boosting、Bootstrap Aggregating和多层融合。According to the human-computer interaction system in the smart home space provided by the present invention, the model fusion methods include but not limited to: Voting, Boosting, Bootstrap Aggregating and multi-layer fusion.
第二方面,本发明提供一种智能家居空间下的人机交互系统的应用方法,所述方法包括:In a second aspect, the present invention provides an application method of a human-computer interaction system in a smart home space, the method comprising:
多通道传感器网络对所述智能家居空间进行多模态信息采集;The multi-channel sensor network collects multi-modal information on the smart home space;
智能感知模块根据采集的多模态信息对所述智能家居空间中的用户进行智能感知;The intelligent perception module performs intelligent perception on users in the smart home space according to the collected multimodal information;
人机交互模块基于用户智能感知结果,实现所述智能家居空间中的各种智能终端设备与所述用户之间的多模态人机交互;The human-computer interaction module realizes multimodal human-computer interaction between various intelligent terminal devices in the smart home space and the user based on the user's intelligent perception results;
其中,所述智能感知至少包括行为感知和情感感知。Wherein, the intelligent perception includes at least behavior perception and emotion perception.
本发明提供一种智能家居空间下的人机交互系统及其应用方法,利用部署于智能家居空间中的多通道传感器网络采集用户智能家居环境下的多源异构数据;基于多源异构数据并利用智能感知模块,感知用户自然行为和情感状态,实现对用户的深度理解;根据用户自然行为和情感状态的感知结果实现智能家居环境下各种智能终端设备与用户之间的多模态人机交互,为用户带来家居环境和家居实体相融合的智慧、舒适、温馨、便捷的个性化家居体验。The present invention provides a human-computer interaction system and its application method in the smart home space, which utilizes the multi-channel sensor network deployed in the smart home space to collect multi-source heterogeneous data in the user smart home environment; based on the multi-source heterogeneous data And use the intelligent perception module to perceive the user's natural behavior and emotional state, and realize the deep understanding of the user; realize the multi-modal human interaction between various intelligent terminal devices and users in the smart home environment according to the perception results of the user's natural behavior and emotional state. Computer interaction, bringing users a personalized home experience that integrates home environment and home entities with wisdom, comfort, warmth and convenience.
附图说明Description of drawings
为了更清楚地说明本发明或现有技术中的技术方案,下面将对实施例或现有技术描述中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图是本发明的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其他的附图。In order to more clearly illustrate the present invention or the technical solutions in the prior art, the accompanying drawings that need to be used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the accompanying drawings in the following description are the For some embodiments of the present invention, those of ordinary skill in the art can also obtain other drawings based on these drawings on the premise of not paying creative efforts.
图1是本发明提供一种智能家居空间下的人机交互系统的结构示意图;Fig. 1 is a schematic structural diagram of a human-computer interaction system under a smart home space provided by the present invention;
图2是本发明提供一种智能家居空间下的人机交互系统的应用方法流程图。Fig. 2 is a flow chart of the application method of the human-computer interaction system in the smart home space provided by the present invention.
具体实施方式Detailed ways
为使本发明的目的、技术方案和优点更加清楚,下面将结合本发明中的附图,对本发明中的技术方案进行清楚、完整地描述,显然,所描述的实施例是本发明一部分实施例,而不是全部的实施例。基于本发明中的实施例,本领域普通技术人员在没有作出创造性劳动前提下所获得的所有其他实施例,都属于本发明保护的范围。In order to make the purpose, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the present invention. Obviously, the described embodiments are part of the embodiments of the present invention , but not all examples. Based on the embodiments of the present invention, all other embodiments obtained by persons of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.
下面结合图1和图2描述本发明提供的一种智能家居空间下的人机交互系统及其应用方法。A human-computer interaction system and its application method in the smart home space provided by the present invention will be described below with reference to FIG. 1 and FIG. 2 .
第一方面,在科研探索层面,智能人居是学科高度交叉的综合研究方向,是与人类活动高度相关的科研领域,涉及人机交互、人工智能、计算机视觉、物联网、用户体验、心理学等学科和领域的综合交叉融合,其相关理论与技术创新,对于深度理解人类行为的模式和情感交流方式具有重要作用,符合人工智能“以人为本”的核心诉求。First, at the level of scientific research and exploration, smart living is a comprehensive research direction with a high degree of cross-discipline, and a scientific research field highly related to human activities, involving human-computer interaction, artificial intelligence, computer vision, Internet of Things, user experience, psychology The comprehensive cross-integration of disciplines and fields, and its related theories and technological innovations play an important role in deeply understanding human behavior patterns and emotional communication methods, which is in line with the core appeal of artificial intelligence "people-oriented".
目前,智能家居的人机交互方式,普遍基于单一智能产品及其应用场景,缺乏普适性和整体性。同时通常以提高单一智能产品的易学性、易用性、用户粘度等为目标,往往忽略了用户的行为意愿和情感体验,达不到期望的效果。At present, the human-computer interaction methods of smart homes are generally based on a single smart product and its application scenarios, which lack universality and integrity. At the same time, it usually aims to improve the ease of learning, ease of use, and user stickiness of a single smart product, often ignoring the user's behavioral intentions and emotional experience, and failing to achieve the desired effect.
基于此,本发明提供一种智能家居空间下的人机交互系统,该系统为多模态人机交互,在人机交互层面上更智能、更精准、更自然。如图1所示,所述系统包括:部署于智能家居空间中的多通道传感器网络110、智能感知模块120和人机交互模块130;Based on this, the present invention provides a human-computer interaction system in a smart home space, which is a multi-modal human-computer interaction system, which is more intelligent, more accurate, and more natural in terms of human-computer interaction. As shown in Figure 1, the system includes: a
其中,所述多通道传感器网络110可以用于采集所述智能家居空间中的多模态信息。Wherein, the
具体的,多通道传感器网络110是由部署于智能家居空间各个位置的传感器构成,因为采用的传感器种类各异。因此采集的多模态信息实则为多源异构数据,即取自多个源头且数据格式迥异的数据,包括但不限于视频数据、图片数据、语音数据、步态数据、气味数据、温度数据、光线数据、湿度数据等。Specifically, the
需要注意的是,对于智能家居用户,很难将完整、全套的传感器网络部署到家居环境中,进行规范化、大量集的数据采集工作。而且目前智能家居用户的房屋多为单一级的智能家居系统,主要实现单品功能或智能家居系统的简单联动,无法形成全套级体系的大范围应用,所收集的数据容易出现不准确或不可用的情况。因此有必要通过智能家居实验研究平台设计多通道传感器网络110中传感器部署方案,以实现智能家居设备的优化控制,进而提高智能家居环境下智能家居设备的可用性和实用性。It should be noted that for smart home users, it is difficult to deploy a complete and complete set of sensor networks in the home environment for standardized and large-scale data collection. Moreover, at present, the houses of smart home users are mostly single-level smart home systems, which mainly realize the simple linkage of single product functions or smart home systems, and cannot form a large-scale application of a full set of level systems, and the collected data is prone to inaccuracy or unavailability Case. Therefore, it is necessary to design a sensor deployment scheme in the
所述智能感知模块120可以用于根据采集的多模态信息对所述智能家居空间中的用户进行智能感知;所述智能感知至少包括行为感知和情感感知;The
具体的,智能感知的概念为通过各种传感器网络获得用户的数据,利用大数据、深度学习、计算机视觉等技术,模仿人和动物的认知机理,完成对象的特征提取和智能推理的过程。本发明除行为感知和情感感知之外,智能感知还可以包括事件感知和环境感知。事件感知,主要针对人与周边物体之间的事件关联。环境感知,主要是指智能家居空间物体所营造的环境的感知。Specifically, the concept of intelligent perception is to obtain user data through various sensor networks, use big data, deep learning, computer vision and other technologies to imitate the cognitive mechanism of humans and animals, and complete the process of object feature extraction and intelligent reasoning. In the present invention, in addition to behavior perception and emotion perception, intelligent perception may also include event perception and environment perception. Event perception is mainly aimed at the event correlation between people and surrounding objects. Environmental perception mainly refers to the perception of the environment created by smart home space objects.
所述智能感知模块120将多模态信息进行脱敏、清洗、标注、特征提取、融合等处理得到表征智能家居环境下用户自然行为和情感状态的多模态特征信息。进而以多模态特征信息为支撑,对用户自然行为和情感状态进行深度理解。The
所述人机交互模块130,用于基于用户智能感知结果,实现所述智能家居空间中的各种智能终端设备与所述用户之间的多模态人机交互;The human-
需要注意的是,本发明在智能家居的原型空间中,利用AIoT、Zigbee等技术实现不同智能终端设备的互融互通、数据共享。It should be noted that in the prototype space of the smart home, the present invention uses technologies such as AIoT and Zigbee to realize the intercommunication and data sharing of different smart terminal devices.
可以理解的是,随着5G技术的正式发牌,智慧物联网时代到来,AI+IoT(即AIoT)已经成为了各大产业巨头的主要赛道,智能家居也正式进入了3.0时代,智能家居即将构建云端互联网平台,实现真正的智能交互、智能感知和复杂决策。在智能家居3.0时代,数据是核心,交互是关键。本发明智能家居空间下的人机交互系统作为3.0时代的一种智能产品,广泛依赖于用户大数据进行深度学习和数据挖掘,精准的刻画用户画像,提供了智能化、个性化的服务。同时还将提供了基于情感计算、语音识别、计算机视觉、知识图谱、自然语言处理等方法的多模态人机交互,实现了“人机物”的深度融合。It is understandable that with the official licensing of 5G technology and the arrival of the era of smart Internet of Things, AI+IoT (ie AIoT) has become the main track of major industry giants, and smart home has officially entered the 3.0 era. It is about to build a cloud Internet platform to realize real intelligent interaction, intelligent perception and complex decision-making. In the era of smart home 3.0, data is the core and interaction is the key. As a smart product in the era of 3.0, the human-computer interaction system in the smart home space of the present invention widely relies on user big data for in-depth learning and data mining, accurately portrays user portraits, and provides intelligent and personalized services. At the same time, it will also provide multi-modal human-computer interaction based on emotional computing, speech recognition, computer vision, knowledge graph, natural language processing and other methods, realizing the deep integration of "human-machine-object".
本发明提供一种智能家居空间下的人机交互系统,利用部署于智能家居空间中的多通道传感器网络110采集用户智能家居环境下的多源异构数据;基于多源异构数据并利用智能感知模块120,感知用户自然行为和情感状态,实现对用户的深度理解;根据用户自然行为和情感状态的感知结果实现智能家居环境下各种智能终端设备与用户之间的多模态人机交互,为用户带来家居环境和家居实体相融合的智慧、舒适、温馨、便捷的个性化家居体验。The present invention provides a human-computer interaction system in the smart home space, which uses the
在上述各实施例的基础上,作为一种可选的实施例,所述多模态信息包括图像信息、语音信息、触觉感知信息、气味信息和环境信息;所述多通道传感器网络110包括:视觉通道、听觉通道、触觉通道、嗅觉通道和环境信息通道;On the basis of the above embodiments, as an optional embodiment, the multimodal information includes image information, voice information, tactile perception information, odor information and environmental information; the
所述视觉通道包括视觉传感器,用于利用所述视觉传感器采集所述图像信息;The visual channel includes a visual sensor for collecting the image information using the visual sensor;
所述听觉通道包括听觉传感器,用于利用所述听觉传感器采集所述语音信息;The auditory channel includes an auditory sensor for collecting the voice information by using the auditory sensor;
所述触觉通道包括触觉传感器,用于利用所述触觉传感器采集所述触觉感知信息;The tactile channel includes a tactile sensor for collecting the tactile perception information by using the tactile sensor;
所述嗅觉通道包括嗅觉传感器,用于利用所述嗅觉传感器采集所述气味信息;The olfactory channel includes an olfactory sensor for collecting the odor information by using the olfactory sensor;
所述环境信息通道包括环境探测传感器,用于利用所述环境探测传感器采集所述环境信息。The environment information channel includes an environment detection sensor for collecting the environment information by using the environment detection sensor.
本发明以面向智能家居的多模态自然人机交互研究为导向,构建以视觉、听觉、触觉、嗅觉为主要模态、环境信息为辅助模态的多通道传感层。在智能家居空间的基础建设节点搭建包括RGB摄像头、高速摄像头、红外摄像头、深度摄像头、麦克风阵列、触觉传感器、步态传感地板、嗅觉传感器、环境信息传感器等硬件的富集传感空间,为多源异构数据的采集提供实现基础。The present invention is oriented towards multi-modal natural human-computer interaction research for smart homes, and constructs a multi-channel sensing layer with vision, hearing, touch, and smell as the main modes and environmental information as the auxiliary mode. Build an enriched sensing space including RGB camera, high-speed camera, infrared camera, depth camera, microphone array, tactile sensor, gait sensing floor, olfactory sensor, environmental information sensor and other hardware at the infrastructure node of the smart home space. The collection of multi-source heterogeneous data provides the basis for realization.
在上述各实施例的基础上,作为一种可选的实施例,所述视觉传感器部署于智能家居空间的房顶,至少包括以下一种:On the basis of the above embodiments, as an optional embodiment, the visual sensor is deployed on the roof of the smart home space, at least including the following one:
RGB相机矩阵、360度全景相机、红外相机和深度相机;RGB camera matrix, 360-degree panoramic camera, infrared camera and depth camera;
所述触觉传感器至少包括以下一种:The tactile sensor includes at least one of the following:
电容地板传感器、布置于智能沙发和/或智能床内的压力传感器;Capacitive floor sensors, pressure sensors arranged in smart sofas and/or smart beds;
所述听觉传感器包括:麦克风阵列;The auditory sensor includes: a microphone array;
所述环境信息传感器至少包括以下一种:The environmental information sensor includes at least one of the following:
温度传感器、湿度传感器、布置于智能床内的二氧化碳传感器和空气质量传感器;Temperature sensor, humidity sensor, carbon dioxide sensor and air quality sensor arranged in the smart bed;
其中,所述电容地板传感器中每个地板面片的电容值由其承受的压力决定。Wherein, the capacitance value of each floor patch in the capacitive floor sensor is determined by the pressure it bears.
具体的,空气质量传感器从功能上分为多种,包括检测PM2.5的空气质量传感器、检测PM10的空气质量传感器、检测二氧化硫的空气质量传感器、检测一氧化碳的空气质量传感器以及其它的空气质量传感器;Specifically, air quality sensors can be divided into multiple types in terms of functions, including air quality sensors that detect PM2.5, air quality sensors that detect PM10, air quality sensors that detect sulfur dioxide, air quality sensors that detect carbon monoxide, and other air quality sensors. ;
视觉通道中布置于智能家居空间的所有相机组成多相机矩阵;RGB相机矩阵指的是由单个视野范围小于360°的RGB相机组成的相机矩阵,从其拍摄的图像或者视频流中可以提取出人体姿态、动作、面部数据、空间位置信息等。360度全景相机指的是360°RGB相机,从其拍摄的图像或者视频流中也可以提取出人体姿态、动作、面部等数据。从红外相机拍摄的图像或者视频流中可以提取出夜间数据、热成像数据等;从深度相机拍摄的图像或者视频流中可以提取出人、空间、物体的深度数据等;All the cameras arranged in the smart home space in the visual channel form a multi-camera matrix; the RGB camera matrix refers to a camera matrix composed of a single RGB camera with a field of view less than 360°, and the human body can be extracted from the captured image or video stream. Posture, action, facial data, spatial location information, etc. A 360-degree panoramic camera refers to a 360-degree RGB camera, and data such as human posture, movement, and face can also be extracted from images or video streams captured by it. Nighttime data, thermal imaging data, etc. can be extracted from images or video streams captured by infrared cameras; depth data of people, space, and objects can be extracted from images or video streams captured by depth cameras;
听觉通道主要由麦克风构成,例如多向麦克风阵列;从多向麦克风阵列中可以提取对话数据。此外由于单个麦克风与音源的距离不同以及所处位置不同,相同声音的声音信号,其波形与频谱信息在不同麦克风中是不同的。因此可以通过这种区别可以从麦克风阵列采集的音频中空间特征信息等。The auditory channel mainly consists of microphones, such as multi-directional microphone arrays; dialogue data can be extracted from multi-directional microphone arrays. In addition, due to the different distances and positions between a single microphone and the sound source, the waveform and spectrum information of the sound signal of the same sound are different in different microphones. Therefore, the spatial feature information and the like in the audio collected by the microphone array can be obtained through this distinction.
触觉通道主要由自研智能家居构成;例如全屋电容地板,布置有压力传感器的智能沙发和智能床;从全屋电容地板中每一片地板的电容值中可以分析出用户位置、步态、方向、速度等用户轨迹信息,可实现活动区域热力图绘制、活动信息记录、多人身份识别和跌倒报警。从智能沙发和智能床感知的压力信息可以分析出用户位置、坐姿、睡眠状态、桌面活动等。The tactile channel is mainly composed of self-developed smart homes; for example, the whole house capacitive floor, smart sofas and smart beds with pressure sensors; the user's position, gait, and direction can be analyzed from the capacitance value of each floor in the whole house capacitive floor , speed and other user trajectory information, which can realize the drawing of heat map of activity area, recording of activity information, multi-person identification and fall alarm. From the pressure information sensed by the smart sofa and smart bed, the user's position, sitting posture, sleep state, desktop activity, etc. can be analyzed.
嗅觉通道主要由嗅觉感知器构成;从嗅觉感知器感知的气味信息中可以识别室内用户活动;The olfactory channel is mainly composed of olfactory sensors; indoor user activities can be identified from the smell information perceived by the olfactory sensors;
环境信息通道主要由各种环境信息传感器构成;从环境信息传感器感知的环境信息中可以获取环境数据,进而识别室内环境状态。The environmental information channel is mainly composed of various environmental information sensors; the environmental data can be obtained from the environmental information sensed by the environmental information sensors, and then the indoor environmental status can be identified.
在本发明给出了视觉传感器、听觉传感器、触觉传感器的示例,为多通道传感器网络110中传感器的选择提供了方向。另外本发明提供的示例并不限制多通道传感器网络110中视觉传感器、听觉传感器、触觉传感器以及嗅觉传感器的选择。Examples of visual sensors, auditory sensors, and tactile sensors are given in the present invention, providing directions for the selection of sensors in the
在上述各实施例的基础上,作为一种可选的实施例,所述多通道传感器网络110,还包括:触发设备;On the basis of the foregoing embodiments, as an optional embodiment, the
所述触发设备,用于以一定的频率触发所述视觉通道、所述听觉通道、所述触觉通道、所述嗅觉通道和所述环境信息通道进行信息的同步采集。The triggering device is configured to trigger the visual channel, the auditory channel, the tactile channel, the olfactory channel and the environmental information channel at a certain frequency to perform synchronous collection of information.
这里,同步采集涵盖采集起停时间以及采集频率的同步。Here, the synchronous acquisition covers the synchronization of acquisition start and stop time and acquisition frequency.
为便于多模态特征信息的融合处理,本发明设置了多通道传感器的集成控制系统,即触发设备,以实现多源异构数据的同步采集。In order to facilitate the fusion processing of multi-modal feature information, the present invention sets up an integrated control system of multi-channel sensors, that is, a trigger device, to realize synchronous collection of multi-source heterogeneous data.
在上述各实施例的基础上,作为一种可选的实施例,所述智能感知模块120包括:预处理单元、多模态特征提取单元和多模态特征融合解析单元;On the basis of the above embodiments, as an optional embodiment, the
所述预处理单元,用于对所述图像信息、所述语音信息、所述触觉感知信息、所述气味信息和所述环境信息进行预处理;The preprocessing unit is configured to preprocess the image information, the voice information, the tactile perception information, the smell information and the environmental information;
所述多模态特征提取单元,用于对预处理后的图像信息、语音信息、触觉感知信息、气味信息和环境信息进行特征提取,得到多模态特征信息;The multimodal feature extraction unit is used to perform feature extraction on the preprocessed image information, voice information, tactile perception information, odor information and environmental information to obtain multimodal feature information;
所述多模态特征融合解析单元,用于对所述多模态特征信息进行融合分析,确定所述用户智能感知结果。The multimodal feature fusion analysis unit is configured to perform fusion analysis on the multimodal feature information to determine the user's intelligent perception result.
在本发明中,采集的所述图像信息、所述语音信息、所述触觉感知信息、所述气味信息和所述环境信息并不能直接使用,需要进行一系列预处理才可以投入使用;示例性的,需要对所述图像信息、所述语音信息、所述触觉感知信息、所述气味信息和所述环境信息进行脱敏,以免泄露用户个人信息;还需要对所述图像信息、所述语音信息、所述触觉感知信息、所述气味信息和所述环境信息进行清洗,归一化等处理,以去除杂乱无用信息。In the present invention, the image information, the voice information, the tactile perception information, the smell information and the environmental information collected cannot be used directly, and a series of preprocessing is required before they can be put into use; exemplary It is necessary to desensitize the image information, the voice information, the tactile perception information, the smell information and the environmental information, so as not to leak the user’s personal information; it is also necessary to desensitize the image information, the voice information Information, the tactile perception information, the odor information and the environmental information are cleaned, normalized, etc., to remove messy and useless information.
在预处理之后,需要对多模态信息进行特征融合分析,以实现对用户的深度智能感知。其中,对多模态信息进行特征融合分析有两种方式:After preprocessing, it is necessary to perform feature fusion analysis on multimodal information to achieve deep intelligent perception of users. Among them, there are two ways to perform feature fusion analysis on multimodal information:
第一种,将多通道数据(即多模态信息,每一个通道来源代表一个模态)进行合并,小尺寸数据使用插值方法(最近邻插值、线性插值、多项插值等)与大尺度数据进行对齐;然后将合并后的多通道数据输入到智能感知模型,得到用户融合分析结果。这种方式中,智能感知模型是利用带有时间戳和智能感知标签的样本数据训练的,其中样本数据包括:对过往多通道数据进行合并后得到的数据。The first one is to combine multi-channel data (that is, multi-modal information, each channel source represents a modality), and use interpolation methods (nearest neighbor interpolation, linear interpolation, multiple interpolation, etc.) for small-scale data and large-scale data Alignment; then input the merged multi-channel data into the intelligent perception model to obtain the user fusion analysis results. In this manner, the intellisense model is trained using sample data with time stamps and intellisense labels, where the sample data includes: data obtained by merging past multi-channel data.
第二种,分别对多模态信息进行特征提取,得到多没模态特征信息;将多没模态特征信息;输入到智能感知模型,得到用户融合分析结果。The second is to extract the features of the multi-modal information separately to obtain the multi-modal feature information; input the multi-modal feature information into the intelligent perception model to obtain the user fusion analysis results.
这种方式中,智能感知模型是将各模态下的智能感知模型进行融合后得到的;其中各模态下的智能感知模型,是利用带有智能感知标签的过往各模型特征信息数据训练的。相对而言,第二种方式比第一种方式的智能感知效果好,因此下述主要对第二种方式进行描述。In this way, the intelligent perception model is obtained by fusing the intelligent perception models in each mode; the intelligent perception model in each mode is trained by using the feature information data of the previous models with the intelligent perception label . Relatively speaking, the Intellisense effect of the second method is better than that of the first method, so the following mainly describes the second method.
本发明融合听觉、视觉、触觉、嗅觉模态(主要模态)和环境信息(辅助模态),构建多元信息和多模态交互的耦合模型,进而全方面的对用户进行智能感知,为智能家居场景下的人机自然交互奠定基础。The present invention integrates auditory, visual, tactile, olfactory modalities (main modal) and environmental information (auxiliary modal), constructs a coupling model of multi-component information and multi-modal interaction, and then intelligently perceives users in all aspects, providing intelligent Lay the foundation for natural human-computer interaction in the home scene.
在上述各实施例的基础上,作为一种可选的实施例,对所述图像信息进行预处理,至少包括以下一种:On the basis of the foregoing embodiments, as an optional embodiment, preprocessing the image information includes at least one of the following:
对所述图像信息进行数据对齐;对所述图像信息进行去冗余;对所述图像信息进行除噪;performing data alignment on the image information; performing de-redundancy on the image information; denoising the image information;
对所述语音信息进行预处理,至少包括以下一种:Preprocessing the voice information includes at least one of the following:
对所述语音信息进行背景音和人声分离;去除所述语音信息中的噪声;去除所述语音信息中的空白音;performing background sound and human voice separation on the voice information; removing noise in the voice information; removing blank tones in the voice information;
对所述触觉感知信息进行预处理,至少包括以下一种:Preprocessing the tactile perception information includes at least one of the following:
对所述触觉感知信息中明显错误数据进行修正;Correcting the obviously wrong data in the tactile perception information;
对所述触觉感知信息中缺失数据进行补全;Completing missing data in the tactile perception information;
对所述气味信息进行预处理,至少包括以下一种:Preprocessing the odor information includes at least one of the following:
去除所述气味信息中包含的辨识失败信息;removing the identification failure information contained in the odor information;
去除所述气味信息中与所述用户的行为习惯明显背离的信息;removing information that obviously deviates from the user's behavior habits in the smell information;
多所述环境信息进行预处理,至少包括以下一种:Preprocessing the environmental information includes at least one of the following:
对所述环境信息中的缺失数据进行补全;Complete missing data in the environmental information;
对所述环境信息中因采集频率所丢失的数据使用插值的方法进行补全。The interpolation method is used to supplement the data lost due to the collection frequency in the environmental information.
本发明中,触觉感知信息一般是压力值、电容值等数值信息,因此在采集过程中会有缺失以及数据明显错的情况,因此对应的预处理措施为不全确实数据以及修正明显错误数据。辨识失败信息没有辨识出气味的信息,即嗅觉传感器辨识失败。与所述用户的行为习惯明显背离的气味信息,例如用户讨厌榴莲气味,如嗅觉传感器辨识榴莲气味大概率是辨识错误。因此对应的预处理措施为去除。In the present invention, the tactile perception information is generally numerical information such as pressure value and capacitance value, so there will be missing and obviously wrong data during the collection process, so the corresponding preprocessing measures are incomplete and correct data and correction of obviously wrong data. The identification failure information does not identify the information of the smell, that is, the identification of the smell sensor fails. Odor information that deviates significantly from the user's behavior habits, for example, the user dislikes the smell of durian, and if the smell sensor recognizes the smell of durian, there is a high probability of misidentification. Therefore, the corresponding pretreatment measure is removal.
可以理解的是,本发明仅示例性的列举了图像信息、语音信息、触觉感知信息、气味信息和环境信息的预处理方式,实际上对图像信息、音频信息、触觉感知信息、气味信息和环境信息的预处理手段不胜枚举,因此可以适应性调整。It can be understood that the present invention only exemplifies the preprocessing methods of image information, voice information, tactile perception information, odor information and environmental information, but in fact, image information, audio information, tactile perception information, odor information and environmental information There are countless means of preprocessing information, so it can be adaptively adjusted.
本发明通过预处理方式,精简多元信息中的有效特征,为后续特征提取奠定基础。The present invention simplifies effective features in multivariate information through a preprocessing method, laying a foundation for subsequent feature extraction.
在上述各实施例的基础上,作为一种可选的实施例,所述多模态特征提取单元,包括:姿态特征提取子单元、人脸特征提取子单元、热成像特征提取子单元、深度特征提取子单元、语义特征提取子单元、位姿特征提取子单元、用户睡眠质量特征提取子单元、用户运动轨迹特征提取子单元、用户行为事件特征提取子单元、温度特征提取子单元、湿度特征提取子单元和空气质量特征提取子单元;On the basis of the above embodiments, as an optional embodiment, the multi-modal feature extraction unit includes: gesture feature extraction subunit, face feature extraction subunit, thermal imaging feature extraction subunit, depth Feature extraction subunit, semantic feature extraction subunit, pose feature extraction subunit, user sleep quality feature extraction subunit, user motion trajectory feature extraction subunit, user behavior event feature extraction subunit, temperature feature extraction subunit, humidity feature extraction subunit and air quality feature extraction subunit;
其中,所述姿态特征提取子单元,用于将所述RGB相机矩阵、所述360度全景相机、处于黑暗环境下的红外相机和/或所述深度相机采集的图像信息经预处理后的数据作为第一目标数据,并基于姿态识别算法从所述第一目标数据中提取姿态特征;Wherein, the posture feature extraction subunit is used to preprocess the image information collected by the RGB camera matrix, the 360-degree panoramic camera, the infrared camera in a dark environment, and/or the depth camera as the first target data, and extract gesture features from the first target data based on a gesture recognition algorithm;
所述人脸特征提取子单元,用于将所述RGB相机矩阵以及所述360度全景相机采集的图像信息经预处理后的数据作为第二目标数据,并基于人脸识别算法从所述第二目标数据中提取人脸特征;The face feature extraction subunit is used to use the preprocessed data of the RGB camera matrix and the image information collected by the 360-degree panoramic camera as the second target data, and obtain the second target data based on the face recognition algorithm. Extract facial features from the target data;
所述热成像特征提取子单元,用于将所述红外相机采集的图像信息经预处理后的数据作为第三目标数据,并从所述第三目标数据中提取人体热成像特征以及物体热成像特征;The thermal imaging feature extraction subunit is used to use the preprocessed data of the image information collected by the infrared camera as the third target data, and extract the thermal imaging features of the human body and the thermal imaging of the object from the third target data feature;
所述深度特征提取子单元,用于将所述深度相机采集的图像信息经预处理后的数据作为第四目标数据,并利用深度学习算法或三维点云特征点提取算法从所述第四目标数据中提取人体深度特征、空间深度特征和物体深度特征;The depth feature extraction subunit is used to use the preprocessed data of the image information collected by the depth camera as the fourth target data, and use a deep learning algorithm or a three-dimensional point cloud feature point extraction algorithm to extract the data from the fourth target Extract human body depth features, spatial depth features and object depth features from the data;
视觉通道使用多相机阵列采集图像信息,采集的图像信息以时间戳为单位并按照图片/视频方式存储。在收集到图片/视频数据后,还会对单帧数据进行图像识别、人体关键骨骼点识别等操作,以从单帧图像中提取,特征包括但不限于:物体信息、空间分布信息等;从多帧人体关键骨骼点信息进行特征提取,所提取特征包括但不限于:肢体动作、动作加速度和位移等。The visual channel uses a multi-camera array to collect image information, and the collected image information is stored in the form of time stamps and pictures/videos. After the picture/video data is collected, operations such as image recognition and key bone point recognition of the human body will be performed on the single-frame data to extract from the single-frame image. Features include but are not limited to: object information, spatial distribution information, etc.; Multi-frame key bone point information of the human body is used for feature extraction. The extracted features include but not limited to: body movements, movement acceleration and displacement, etc.
所述语义特征提取子单元,用于将所述麦克风阵列采集的语音信息经预处理后的数据作为第五目标数据,并利用语音识别算法和/或语音情感分析算法,从所述第五目标数据中提取语义特征;The semantic feature extraction subunit is configured to use the preprocessed data of the voice information collected by the microphone array as the fifth target data, and use a voice recognition algorithm and/or a voice emotion analysis algorithm to extract the data from the fifth target Extract semantic features from the data;
听觉通道使用麦克风阵列按一定采样率收集音频信息,音频信息以时间戳为单位存储。对音频信息进行特征提取,提取特征的方法包括但不限于傅里叶变换、频谱图、奈奎斯特采样,提取过零率、频谱中心、频谱滚降点、梅尔频率倒谱系数(MFCC)等。The auditory channel uses a microphone array to collect audio information at a certain sampling rate, and the audio information is stored in units of time stamps. Feature extraction of audio information, the method of feature extraction includes but not limited to Fourier transform, spectrogram, Nyquist sampling, extraction of zero-crossing rate, spectral center, spectral roll-off point, Mel frequency cepstral coefficient (MFCC )Wait.
所述位姿特征提取子单元,用于将所述布置于智能沙发和/或智能床内的压力传感器采集的压力值信息经预处理后的数据作为第六目标数据,并分析所述第六目标数据得到用户在所述智能沙发和/或智能床上的位姿特征;The pose feature extraction subunit is used to use the preprocessed data of the pressure value information collected by the pressure sensor arranged in the smart sofa and/or smart bed as the sixth target data, and analyze the sixth target data. The target data obtains the pose characteristics of the user on the smart sofa and/or smart bed;
进一步的,从用户睡眠质量特征还可以衍生出用户睡眠健康等维度特征。Furthermore, dimensional features such as the user's sleep health can also be derived from the user's sleep quality feature.
所述用户运动轨迹特征提取子单元,用于将所述电容地板传感器采集的电容值信息经预处理后的数据作为第七目标数据,并分析所述第七目标数据得到用户运动轨迹特征;所述用户运动轨迹特征包括:位置、步态、方向和速度;The user movement trajectory feature extraction subunit is used to use the preprocessed data of the capacitance value information collected by the capacitive floor sensor as the seventh target data, and analyze the seventh target data to obtain the user movement trajectory characteristics; The characteristics of the user's motion trajectory include: position, gait, direction and speed;
所述用户行为事件特征提取子单元,用于将所述嗅觉传感器采集的气味信息经预处理后的数据作为第八目标数据,并分析所述第八目标数据得到用户行为事件特征。The user behavior event feature extraction subunit is configured to use the preprocessed data of the smell information collected by the olfactory sensor as the eighth target data, and analyze the eighth target data to obtain the user behavior event features.
具体的,用户行为事件特征指的是不同空间正在发生的人类日常生活行为特征。Specifically, user behavior event features refer to the characteristics of human daily life behaviors that are happening in different spaces.
所述用户睡眠质量特征子提取单元,用于将所述布置于智能床内的二氧化碳传感器采集的二氧化碳值信息经预处理后的数据作为第九目标数据,并分析所述第九目标数据得到用户睡眠质量特征;The user sleep quality feature sub-extraction unit is used to use the preprocessed data of the carbon dioxide value information collected by the carbon dioxide sensor arranged in the smart bed as the ninth target data, and analyze the ninth target data to obtain the user sleep quality characteristics;
所述温度特征提取子单元,用于将所述温度传感器采集的温度信息经预处理后的数据作为第十目标数据,并分析所述第十目标数据得到温度特征;The temperature feature extraction subunit is configured to use the preprocessed data of temperature information collected by the temperature sensor as the tenth target data, and analyze the tenth target data to obtain a temperature feature;
所述湿度特征提取子单元,用于将所述湿度传感器采集的湿度信息经预处理后的数据作为第十一目标数据,并分析所述第十一目标数据得到湿度特征;The humidity feature extraction subunit is configured to use the preprocessed data of the humidity information collected by the humidity sensor as the eleventh target data, and analyze the eleventh target data to obtain a humidity feature;
所述空气质量特征提取子单元,用于将所述空气质量传感器采集的空气质量信息经预处理后的数据作为第十二目标数据,并分析所述第十二目标数据得到空气质量特征。The air quality feature extraction subunit is configured to use the preprocessed data of the air quality information collected by the air quality sensor as the twelfth target data, and analyze the twelfth target data to obtain the air quality feature.
本发明给出了多种通道采集的数据的特征提取方式,为多模态特征数据的分析学习奠定基础。The invention provides the feature extraction mode of data collected by multiple channels, and lays the foundation for the analysis and learning of multi-modal feature data.
在上述各实施例的基础上,作为一种可选的实施例,所述多模态特征融合解析单元,包括:数据存储单元、智能感知模型构造单元和智能感知执行单元;On the basis of the above embodiments, as an optional embodiment, the multi-modal feature fusion analysis unit includes: a data storage unit, an intelligent perception model construction unit, and an intelligent perception execution unit;
所述数据存储单元,用于将所述多模态特征中各个模态特征存储于各个模态特征数据库中;The data storage unit is used to store each modal feature in the multi-modal feature in each modal feature database;
具体的,所述智能感知模块120从多模态信息中提取表征智能家居环境下用户自然行为和情感状态的多模态特征信息。由于多通道传感器网络110采集的多模态信息具备多维度、价值密度低、产生速度快等特点,因此可以在短期内构建多模态特征数据库,为基于智能感知的多模态人机交互提供数据支撑。Specifically, the
所述智能感知模型构造单元,用于在训练阶段利用各个模态特征数据库训练基于各个模态特征的智能感知模型;并将所述基于各个模态特征的智能感知模型进行模型融合得到智能感知模型;还用于在应用阶段利用所述多模态特征更新所述智能感知模型;The intelligent perception model construction unit is used to train the intelligent perception model based on each modal feature using each modal feature database in the training phase; and the intelligent perception model based on each modal feature is model-fused to obtain the intelligent perception model ; Also used to update the intellisense model using the multi-modal feature in the application phase;
所述智能感知执行单元,用于将所述多模态特征输入所述智能感知模型,得到所述智能感知模型输出的用户智能感知结果。The intelligent perception executing unit is configured to input the multimodal features into the intelligent perception model, and obtain the user's intelligent perception result output by the intelligent perception model.
本发明将各类通道的特征数据,输入到面向智能家居的多模态多目标识别算法模型中,通过多模态特征数据的分析学习,输出智能家居环境下用户行为、用户情感、发生事件、环境状态等,用于指导智能家居环境下的各种智能终端设备(例如家庭服务器人、智能家电等)与用户开展多模态人机交互。In the present invention, the feature data of various channels are input into the multi-modal multi-target recognition algorithm model for smart home, and through the analysis and learning of multi-modal feature data, user behavior, user emotion, occurrence events, Environmental status, etc., are used to guide various smart terminal devices (such as home servers, smart home appliances, etc.) in the smart home environment to carry out multi-modal human-computer interaction with users.
在上述各实施例的基础上,作为一种可选的实施例,所述模型融合方式,包括但不限于:Voting、Boosting、Bootstrap Aggregating和多层融合。On the basis of the foregoing embodiments, as an optional embodiment, the model fusion manner includes, but is not limited to: Voting, Boosting, Bootstrap Aggregating and multi-layer fusion.
本发明给出了模型融合方式的选择方向,且并未限制模型融合方式的选择范围。The present invention provides the selection direction of the model fusion mode, and does not limit the selection range of the model fusion mode.
第二方面,本发明提供一种智能家居空间下的人机交互系统的应用方法,如图2所示,所述方法包括:In a second aspect, the present invention provides an application method of a human-computer interaction system in a smart home space, as shown in FIG. 2 , the method includes:
S11、多通道传感器网络对所述智能家居空间进行多模态信息采集;S11. The multi-channel sensor network collects multi-modal information on the smart home space;
S12、智能感知模块根据采集的多模态信息对所述智能家居空间中的用户进行智能感知;S12. The intelligent perception module performs intelligent perception on users in the smart home space according to the collected multimodal information;
S13、人机交互模块基于用户智能感知结果,实现所述智能家居空间中的各种智能终端设备与所述用户之间的多模态人机交互;S13. The human-computer interaction module implements multimodal human-computer interaction between various intelligent terminal devices in the smart home space and the user based on the user's intelligent perception result;
其中,所述智能感知至少包括行为感知和情感感知。Wherein, the intelligent perception includes at least behavior perception and emotion perception.
本发明提供一种智能家居空间下的人机交互系统的应用方法,利用部署于智能家居空间中的多通道传感器网络采集用户智能家居环境下的多源异构数据;基于多源异构数据并利用智能感知模块,感知用户自然行为和情感状态,实现对用户的深度理解;根据用户自然行为和情感状态的感知结果实现智能家居环境下各种智能终端设备与用户之间的多模态人机交互,为用户带来家居环境和家居实体相融合的智慧、舒适、温馨、便捷的个性化家居体验。The invention provides an application method of a human-computer interaction system in a smart home space, which uses a multi-channel sensor network deployed in a smart home space to collect multi-source heterogeneous data in a user's smart home environment; based on multi-source heterogeneous data and Use the intelligent perception module to perceive the user's natural behavior and emotional state, and realize a deep understanding of the user; realize the multi-modal human-machine interaction between various smart terminal devices and users in the smart home environment based on the perception results of the user's natural behavior and emotional state Interaction brings users a personalized home experience that integrates home environment and home entities with wisdom, comfort, warmth and convenience.
以上所描述的装置实施例仅仅是示意性的,其中所述作为分离部件说明的单元可以是或者也可以不是物理上分开的,作为单元显示的部件可以是或者也可以不是物理单元,即可以位于一个地方,或者也可以分布到多个网络单元上。可以根据实际的需要选择其中的部分或者全部模块来实现本实施例方案的目的。本领域普通技术人员在不付出创造性的劳动的情况下,即可以理解并实施。The device embodiments described above are only illustrative, and the units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in One place, or it can be distributed to multiple network elements. Part or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. It can be understood and implemented by those skilled in the art without any creative efforts.
通过以上的实施方式的描述,本领域的技术人员可以清楚地了解到各实施方式可借助软件加必需的通用硬件平台的方式来实现,当然也可以通过硬件。基于这样的理解,上述技术方案本质上或者说对现有技术做出贡献的部分可以以软件产品的形式体现出来,该计算机软件产品可以存储在计算机可读存储介质中,如ROM/RAM、磁碟、光盘等,包括若干指令用以使得一台计算机设备(可以是个人计算机,服务器,或者网络设备等)执行各个实施例或者实施例的某些部分所述的方法。Through the above description of the implementations, those skilled in the art can clearly understand that each implementation can be implemented by means of software plus a necessary general-purpose hardware platform, and of course also by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the prior art can be embodied in the form of software products, and the computer software products can be stored in computer-readable storage media, such as ROM/RAM, magnetic discs, optical discs, etc., including several instructions to make a computer device (which may be a personal computer, server, or network device, etc.) execute the methods described in various embodiments or some parts of the embodiments.
最后应说明的是:以上实施例仅用以说明本发明的技术方案,而非对其限制;尽管参照前述实施例对本发明进行了详细的说明,本领域的普通技术人员应当理解:其依然可以对前述各实施例所记载的技术方案进行修改,或者对其中部分技术特征进行等同替换;而这些修改或者替换,并不使相应技术方案的本质脱离本发明各实施例技术方案的精神和范围。Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: it can still be Modifications are made to the technical solutions described in the foregoing embodiments, or equivalent replacements are made to some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims (10)
Priority Applications (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210861645.9A CN115291718A (en) | 2022-07-20 | 2022-07-20 | Man-machine interaction system in smart home space and application method thereof |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202210861645.9A CN115291718A (en) | 2022-07-20 | 2022-07-20 | Man-machine interaction system in smart home space and application method thereof |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| CN115291718A true CN115291718A (en) | 2022-11-04 |
Family
ID=83823764
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| CN202210861645.9A Pending CN115291718A (en) | 2022-07-20 | 2022-07-20 | Man-machine interaction system in smart home space and application method thereof |
Country Status (1)
| Country | Link |
|---|---|
| CN (1) | CN115291718A (en) |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116578861A (en) * | 2023-04-27 | 2023-08-11 | 青岛海尔科技有限公司 | Equipment state control method and device, storage medium and electronic device |
| CN117152420A (en) * | 2023-10-26 | 2023-12-01 | 享刻智能技术(北京)有限公司 | Perception methods, systems and robots based on multimodal events |
| CN118068736A (en) * | 2024-02-26 | 2024-05-24 | 贝塔智能科技(北京)有限公司 | A spray intelligent sensing control method and system |
| CN120010282A (en) * | 2025-04-21 | 2025-05-16 | 深圳市集贤科技有限公司 | A whole-house intelligent scene generation system based on natural language |
| CN121331126A (en) * | 2025-10-30 | 2026-01-13 | 江西澳客家居科技有限公司 | Sofa Control Method and System Based on Intelligent Voice |
Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104302016A (en) * | 2014-09-16 | 2015-01-21 | 北京市信息技术研究所 | A wireless sensor network architecture based on multifunctional composite sensors |
| CN105827731A (en) * | 2016-05-09 | 2016-08-03 | 包磊 | Intelligent health management server, system and control method based on fusion model |
| CN111538251A (en) * | 2020-05-22 | 2020-08-14 | 江洪华 | Method and system for optimizing environment |
| US20210103762A1 (en) * | 2019-10-02 | 2021-04-08 | King Fahd University Of Petroleum And Minerals | Multi-modal detection engine of sentiment and demographic characteristics for social media videos |
| CN113902965A (en) * | 2021-09-30 | 2022-01-07 | 重庆邮电大学 | Multi-spectral pedestrian detection method based on multi-layer feature fusion |
-
2022
- 2022-07-20 CN CN202210861645.9A patent/CN115291718A/en active Pending
Patent Citations (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN104302016A (en) * | 2014-09-16 | 2015-01-21 | 北京市信息技术研究所 | A wireless sensor network architecture based on multifunctional composite sensors |
| CN105827731A (en) * | 2016-05-09 | 2016-08-03 | 包磊 | Intelligent health management server, system and control method based on fusion model |
| US20210103762A1 (en) * | 2019-10-02 | 2021-04-08 | King Fahd University Of Petroleum And Minerals | Multi-modal detection engine of sentiment and demographic characteristics for social media videos |
| CN111538251A (en) * | 2020-05-22 | 2020-08-14 | 江洪华 | Method and system for optimizing environment |
| CN113902965A (en) * | 2021-09-30 | 2022-01-07 | 重庆邮电大学 | Multi-spectral pedestrian detection method based on multi-layer feature fusion |
Non-Patent Citations (1)
| Title |
|---|
| 付心仪等: "智能家居综合实验平台设计研究与应用实践", 包装工程, vol. 43, no. 16, 31 August 2022 (2022-08-31) * |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| CN116578861A (en) * | 2023-04-27 | 2023-08-11 | 青岛海尔科技有限公司 | Equipment state control method and device, storage medium and electronic device |
| CN117152420A (en) * | 2023-10-26 | 2023-12-01 | 享刻智能技术(北京)有限公司 | Perception methods, systems and robots based on multimodal events |
| CN118068736A (en) * | 2024-02-26 | 2024-05-24 | 贝塔智能科技(北京)有限公司 | A spray intelligent sensing control method and system |
| CN120010282A (en) * | 2025-04-21 | 2025-05-16 | 深圳市集贤科技有限公司 | A whole-house intelligent scene generation system based on natural language |
| CN121331126A (en) * | 2025-10-30 | 2026-01-13 | 江西澳客家居科技有限公司 | Sofa Control Method and System Based on Intelligent Voice |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Cruciani et al. | Feature learning for human activity recognition using convolutional neural networks: A case study for inertial measurement unit and audio data | |
| US11226673B2 (en) | Affective interaction systems, devices, and methods based on affective computing user interface | |
| Schiel et al. | The SmartKom Multimodal Corpus at BAS. | |
| Scherer et al. | A generic framework for the inference of user states in human computer interaction: How patterns of low level behavioral cues support complex user states in HCI | |
| CN110519636A (en) | Voice messaging playback method, device, computer equipment and storage medium | |
| JP2018014094A (en) | Virtual robot interaction method, system, and robot | |
| CN120182488A (en) | An immersive exhibition hall intelligent guided tour display method and system based on user behavior | |
| CN120413044B (en) | Multi-mode data interaction-based pediatric nursing quality assessment method | |
| CN120315314B (en) | Intelligent home scene control method and system based on edge calculation | |
| CN119004195B (en) | Intelligent projection method, system, medium and program product based on gesture recognition | |
| CN119996786A (en) | Video content generation method, system, device and medium for multi-source material fusion | |
| CN120068923A (en) | Digital human intelligent interaction and gesture expression synthesis method based on multi-modal synchronization | |
| Verma et al. | Affective state recognition from hand gestures and facial expressions using Grassmann manifolds | |
| CN120340481B (en) | A display voice interaction system and method | |
| CN121210874A (en) | Sand table intelligent recognition system based on artificial intelligence multimodal technology | |
| CN115268287A (en) | Intelligent home comprehensive experiment system and data processing method | |
| CN120524447A (en) | Emotional interaction method and device based on multimodal data fusion | |
| Monekosso et al. | Intelligent environments: methods, algorithms and applications | |
| KR102752005B1 (en) | Care robot for providing care service | |
| CN117942079B (en) | Emotion intelligence classification method and system based on multidimensional sensing and fusion | |
| CN118196702A (en) | Emotion monitoring method and system for remote video communication personnel based on domain generalization | |
| Jyothsna et al. | Face recognition automated system for visually impaired peoples using machine learning | |
| CN121434508B (en) | Multi-mode intelligent body driven digital librarian interaction optimization method and system | |
| CN117854666B (en) | A method and device for constructing a three-dimensional human rehabilitation data set | |
| CN118963603B (en) | Intelligent panel control method based on DNN noise reduction technology |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| PB01 | Publication | ||
| PB01 | Publication | ||
| SE01 | Entry into force of request for substantive examination | ||
| SE01 | Entry into force of request for substantive examination |
