WO2025201223A1 - 用于智能辅助交互的电子设备和方法 - Google Patents
用于智能辅助交互的电子设备和方法Info
- Publication number
- WO2025201223A1 WO2025201223A1 PCT/CN2025/084290 CN2025084290W WO2025201223A1 WO 2025201223 A1 WO2025201223 A1 WO 2025201223A1 CN 2025084290 W CN2025084290 W CN 2025084290W WO 2025201223 A1 WO2025201223 A1 WO 2025201223A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- mode
- user
- electronic device
- model
- inference
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F11/00—Error detection; Error correction; Monitoring
- G06F11/36—Prevention of errors by analysis, debugging or testing of software
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F18/00—Pattern recognition
- G06F18/20—Analysing
- G06F18/24—Classification techniques
- G06F18/243—Classification techniques relating to the number of classes
- G06F18/2431—Multiple classes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/011—Arrangements for interaction with the human body, e.g. for user immersion in virtual reality
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F3/00—Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
- G06F3/01—Input arrangements or combined input and output arrangements for interaction between user and computer
- G06F3/017—Gesture based interaction, e.g. based on a set of recognized hand gestures
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06N—COMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
- G06N5/00—Computing arrangements using knowledge-based models
- G06N5/04—Inference or reasoning models
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T1/00—General purpose image data processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06T—IMAGE DATA PROCESSING OR GENERATION, IN GENERAL
- G06T1/00—General purpose image data processing
- G06T1/0021—Image watermarking
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/70—Arrangements for image or video recognition or understanding using pattern recognition or machine learning
- G06V10/764—Arrangements for image or video recognition or understanding using pattern recognition or machine learning using classification, e.g. of video objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/40—Scenes; Scene-specific elements in video content
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/50—Context or environment of the image
- G06V20/52—Surveillance or monitoring of activities, e.g. for recognising suspicious objects
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
- G06V40/23—Recognition of whole body movements, e.g. for sport training
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V40/00—Recognition of biometric, human-related or animal-related patterns in image or video data
- G06V40/20—Movements or behaviour, e.g. gesture recognition
- G06V40/28—Recognition of hand or arm movements, e.g. recognition of deaf sign language
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/28—Constructional details of speech recognition systems
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V10/00—Arrangements for image or video recognition or understanding
- G06V10/20—Image preprocessing
- G06V10/26—Segmentation of patterns in the image field; Cutting or merging of image elements to establish the pattern region, e.g. clustering-based techniques; Detection of occlusion
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06V—IMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
- G06V20/00—Scenes; Scene-specific elements
- G06V20/70—Labelling scene content, e.g. deriving syntactic or semantic representations
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/223—Execution procedure of a spoken command
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/225—Feedback of the input speech
-
- G—PHYSICS
- G10—MUSICAL INSTRUMENTS; ACOUSTICS
- G10L—SPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
- G10L15/00—Speech recognition
- G10L15/22—Procedures used during a speech recognition process, e.g. man-machine dialogue
- G10L2015/226—Procedures used during a speech recognition process, e.g. man-machine dialogue using non-speech characteristics
Definitions
- the present disclosure relates to the field of computers, and more particularly, to intelligent auxiliary interaction.
- intelligent interactive devices have gradually become part of our daily lives.
- intelligent voice assistants can receive natural language input and, through AI-based technologies such as speech recognition and semantic analysis, fully understand the user's intent, providing relevant information or continuous interaction.
- the intelligent assisted interaction systems known to the inventors of this disclosure primarily take the form of various voice-activated assistants, such as smart speakers and chatbots, but lack intelligent services that utilize image input.
- interactive systems that utilize image output e.g., monitoring systems
- image output are relatively low in intelligence. For example, they can only detect simple functions such as crying babies or the presence of people, but struggle to meet customized requirements for monitoring complex scenarios (e.g., determining whether the monitored individual is studying or exercising at a specified time).
- the present disclosure provides an electronic device and method for intelligent assisted interaction, which can provide integrated voice and image interaction, and support user-customized complex scenario service settings to enhance the user's interactive experience.
- a computer program product including a computer program, wherein when the computer program is executed by a processor, the processor is caused to execute the intelligent auxiliary interaction method according to the present disclosure.
- FIG4 is a schematic diagram illustrating inference metadata according to an embodiment of the present disclosure.
- FIG5 is an exemplary flowchart illustrating an image encryption process according to an embodiment of the present disclosure
- FIG1 shows an exemplary configuration block diagram of an electronic device 1000 for intelligent auxiliary interaction according to an embodiment of the present disclosure and a system environment in which the electronic device 1000 is applied.
- the intelligent assistance interaction described in this disclosure refers to the use of artificial intelligence technologies to enable interaction between users and intelligent assistance systems (e.g., electronic device 1000), thereby providing users with personalized, intelligent services and assistance. It can support a variety of different interaction methods, including but not limited to voice, image, and gesture, and can utilize multimodal information to interact with the intelligent assistance system.
- the electronic device 1000 is communicatively coupled to the user terminal device 1100, receives the function setting for intelligent assisted interaction from the user terminal device 1100, and determines the operating mode of the electronic device to be used for intelligent assisted interaction based on the function setting, so as to perform the corresponding intelligent assisted interaction action in the operating mode.
- the electronic device 1000 can send the results (such as inference results) obtained by performing the intelligent assisted interaction action to the user terminal device 1100 so that the user can obtain the required information.
- the electronic device 1000 may also be communicatively coupled to the cloud platform 1200 to update the AI model used to perform the corresponding intelligent assistance interaction action from the cloud platform 1200.
- the electronic device 1000 may also send inference results to the cloud platform 1200 and/or receive function settings via the cloud platform 1200.
- the cloud platform 1200 and the user terminal device 1100 may be communicatively coupled.
- the user terminal device 1100 may send function settings for intelligent assisted interaction to the cloud platform 1200, and receive inference results of intelligent assisted interaction actions from the electronic device 1000 via the cloud platform 1200.
- electronic device 1000 may include a processing circuit 1010.
- the processing circuit 1010 of electronic device 1000 provides various functions of electronic device 1000.
- the processing circuit 1010 of electronic device 1000 may be configured to execute a method for electronic device 1000.
- user-end device 1100 may also include a processing circuit (not shown) for providing various functions of user-end device 1100.
- Processing circuitry 1010 may refer to various implementations of digital circuitry, analog circuitry, or mixed-signal (a combination of analog and digital) circuitry that performs functions in a computing system.
- Processing circuitry may include, for example, circuits such as integrated circuits (ICs), application-specific integrated circuits (ASICs), portions or circuits of a separate processor core, an entire processor core, a separate processor, a programmable hardware device such as a field-programmable gate array (FPGA), and/or a system including multiple processors.
- ICs integrated circuits
- ASICs application-specific integrated circuits
- FPGA field-programmable gate array
- the processing circuit 1010 may include an information receiving unit 1020, an inference unit 1030, a model preparation unit 1040 and a post-processing unit, configured to execute corresponding steps in the method 2000 shown in FIG. 2 described below.
- the electronic device 1000 may further include a memory (not shown).
- the memory of the electronic device 1000 may store information generated by the processing circuit 1010 as well as programs and data used for the operation of the electronic device 1000.
- the memory may be a volatile memory and/or a non-volatile memory.
- the memory may include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read-only memory (ROM), and a flash memory.
- the electronic device 1000 may be implemented at a chip level, or may be implemented at a device level by including other external components. In some embodiments, the electronic device 1000 may be implemented as a complete device, and may further include multiple antennas.
- the above-mentioned units are merely logical modules divided according to the specific functions they implement, and are not intended to limit the specific implementation methods.
- the above-mentioned units can be implemented as independent physical entities, or can be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.).
- FIG2 shows an exemplary flowchart of an intelligent auxiliary interaction method 2000 according to an embodiment of the present disclosure.
- the method 2000 may be implemented by the electronic device 1000 shown in FIG1 , for example.
- the information receiving unit 1020 receives the user's function settings for the electronic device 1000.
- the function settings are used, for example, to set the user's desired operating mode for the electronic device 1000 and the method for receiving inference results, so that the electronic device 1000 operates in the operating mode to implement the user's desired auxiliary interaction function and returns the inference results in the method desired by the user.
- the user can download the local client program at the user terminal device 1100 (for example, through his own local terminal device (such as a computer, tablet computer, smart phone, smart watch, etc.)) and perform function settings through the user terminal device 1100.
- his own local terminal device such as a computer, tablet computer, smart phone, smart watch, etc.
- the processing circuit of the user terminal device 1100 can be configured to display an interactive interface for the user to set the functions of the electronic device 1000.
- the interactive interface can display a menu option of optional operating modes for the user to select and set the desired operating mode of the electronic device 1000.
- the operating mode may include a user-customized service mode and a developer debugging mode.
- typical services required by the user may be provided, while in the developer debugging mode, auxiliary tools/information required by the developer for development or debugging may be provided.
- menu options for the user-customized service mode and the developer debugging mode may be displayed in the interactive interface for the user to select as needed.
- a secondary menu in response to the user selecting the user customized service mode, may be further displayed in the interactive interface for the user to select a sub-mode under the user customized service mode, such as game mode, home mode, smart assistant mode, etc.
- a tertiary menu may be further displayed in the interactive interface for the user to select a further subdivided service mode under the sub-mode.
- modes such as yoga practice (monitoring whether yoga movements are correct), singing and dancing practice (monitoring whether singing and dancing practices meet standards and assisting in improvement) may be included; in the home mode, modes such as monitoring student activities (monitoring whether learning is done during the set learning period) and health monitoring (assessing physical health by capturing movements or body shapes) may be included; in the smart assistant mode, modes for persons with disabilities may be included, such as sign language and motion recognition (for example, controlling music playback through sign language) for persons with disabilities (especially deaf-mute persons).
- multiple sub-modes in this mode can be displayed in the secondary menu of the interactive interface.
- the sub-modes may include game developer mode and algorithm developer mode, etc.
- the user may be provided with intermediate results such as human body position information and human motion recognition results to assist the user in game development;
- the algorithm developer mode the user may be provided with human body key point information (such as two-dimensional or three-dimensional key point information), etc., to assist in the development of the corresponding algorithm.
- the user can select the required sub-mode in this secondary menu.
- the sub-modes include classification mode, detection mode, text recognition mode, image segmentation mode, semantic segmentation mode, etc., for developers to select according to their needs.
- the sub-modes can also be further classified for users to make more detailed choices.
- the detection mode can be further classified into conventional detection mode, key point detection mode, face detection mode, etc.
- a user can perform function settings through the user terminal device 1100 to set the operating mode of the electronic device 1000.
- the processing circuit of the user terminal device 1100 can be configured to provide the function settings to the electronic device 1000, and the information receiving unit 1020 of the electronic device 1000 receives the function settings and determines the operating mode of the electronic device 1000 accordingly for subsequent processing.
- multiple AI models including various types of speech processing models and image processing models, may be pre-stored in the memory of the electronic device 1000.
- the model preparation unit 1030 may select the AI model required for the operating mode from the pre-stored multiple AI models.
- a corresponding image processing model (such as a human body key point analysis model, etc.) can be selected from multiple AI models to analyze the user's movements when practicing yoga and perform corresponding reasoning.
- corresponding voice processing models and image processing models can be selected from multiple AI models to analyze and reason about the user's voice and movements, respectively.
- the AI model required by the developer can also be selected from multiple AI models.
- a plurality of pre-prepared AI models may also be stored (or partially stored) in the cloud platform 1200.
- the AI model can be updated using the cloud platform 1200.
- the corresponding AI model can be downloaded from the cloud platform 1200 to the electronic device 1000 for subsequent use.
- the electronic device 1000 can be used offline without relying on network conditions, thereby expanding the scope of use.
- the electronic device 1000 may include a voice/image acquisition unit for acquiring voice and/or image information to be inferred, such as the voice and movements of a user's yoga movements, singing and dancing exercises, etc.
- the electronic device 1000 may also be communicatively connected to a voice/image acquisition device (e.g., a microphone, a camera, etc.) to receive the voice and/or image information acquired by the voice/image acquisition unit.
- a voice/image acquisition device e.g., a microphone, a camera, etc.
- the set operating mode is the yoga practice mode in the user-customized service mode
- the user's yoga movement images/videos can be collected, and the prepared human body key point analysis model can be used to extract key points, thereby inferring the user's yoga movements to detect whether the yoga movements are correct.
- the set operating mode is the singing and dancing practice scenario in the user-customized service mode
- the user's voice and images/videos during singing and dancing practice can be collected, and the corresponding voice processing AI model can be used to infer the practice status of the singing part
- the image processing AI model can be used to infer the practice status of the dancing part, thereby assisting the user's singing and dancing practice.
- the post-processing unit 1050 determines the user's reception mode for the inference result according to the user's function setting, and processes the inference result according to the determined reception mode.
- the processing of the inference result is referred to as "post-processing”.
- an application program interface API related to the result of the reasoning can be selected to be provided to the user, so as to provide the user (developer) with a lower-level API service, so as to facilitate the developer to perform reasoning metadata queries (such as obtaining two-dimensional and three-dimensional key point coordinates through the Get Coordinate Point API, or obtaining the current motion recognition results of the subject through the Get Recognition Result API, etc.), thereby accelerating its development process.
- FIG. 3 an exemplary flowchart of an intelligent auxiliary interaction method 3000 according to another embodiment of the present disclosure is described.
- the method 3000 may also be implemented by the electronic device 1000 shown in FIG. 1 , for example.
- step S3010 it is determined whether the operating mode is set by the user. This step can be determined, for example, by whether function settings are received from the user terminal device 1100 or the cloud platform 1200 , or whether the received function settings include settings regarding the operating mode.
- step S3020 to prepare the corresponding AI model according to the set operating mode. This step corresponds to step S2020 described with reference to FIG.
- a sign language recognition algorithm model can be prepared in advance to identify the sign language in the collected image/video information of the disabled person for use when the operating mode is the default disabled person assistant mode.
- the sign language recognition algorithm model can be pre-stored in the memory of the electronic device 1000, for example, without downloading updates from the cloud platform 1300, thereby supporting offline use of the disabled person assistant mode. In this way, intelligent assisted interaction between disabled people (especially deaf-mute people) and electronic devices can be achieved, improving the user experience.
- a speech recognition algorithm model can also be prepared in advance to identify voice interactions of disabled people such as the blind.
- the "Assistance Mode for Persons with Disabilities" is set as the default operating mode.
- Persons with disabilities such as deaf-mute persons, blind persons, etc.
- those skilled in the art may also set other operating modes as the default mode according to actual needs, or may not set a default operating mode.
- the types of operating modes are not limited to the three types listed above, and may include other operating modes, or classify the operating modes in other ways.
- step S3025 the prepared AI model is used to perform inference on the collected voice and/or image information. This step corresponds to step S2030 described with reference to FIG.
- the voice, movement, gesture, etc. of the monitored object can be collected within a predetermined time period or according to a predetermined cycle for use in reasoning.
- the user when the operating mode is the student activity monitoring mode, the user (e.g., a parent) can further set a monitoring time period through the user terminal device 1100 as a function setting for the electronic device 1000 and provide it to the electronic device 1000.
- the electronic device 1000 controls its image acquisition unit to capture images of the detection object (e.g., student) according to the set monitoring time period, or controls the image acquisition device connected thereto to capture images, thereby monitoring whether the student is studying during the set study period.
- the monitored object can be the same as or different from the user.
- the monitored object and the user are different.
- the user and the monitored object can also be the same.
- Electronic device 1000 captures the user's singing and dancing practice videos within a set time period to monitor whether the singing and dancing practice meets the standards and assist in improvement.
- step S3030 it is determined whether the operating mode is developer mode. In some embodiments, this step can also be combined with the determination process of step S3015, or the determination result of step S3015 can be used. If the determination in step S3030 is "yes”, the process proceeds to step S3040 to provide API services to the user. In addition, if the determination in step S3030 is "no", the result of the inference obtained in step S3025 is converted into readable information and provided to the user.
- the user's reception mode for the result of the inference can be further determined.
- the reception mode can be determined based on the operating mode set by the user. In other embodiments, the user can also set the reception mode at the user terminal device 1100.
- Computing device 600 is an example of a hardware device to which the above-described aspects of the present invention can be applied.
- Computing device 600 can be any machine configured to perform processing and/or computations.
- Computing device 600 can be, but is not limited to, a workstation, a server, a desktop computer, a laptop computer, a tablet computer, a personal data assistant (PDA), a smartphone, an in-vehicle computer, or a combination thereof.
- PDA personal data assistant
- the computing device 600 may also include or be connected to a non-transitory storage device 614, which may be any non-transitory storage device that can implement data storage and may include, but is not limited to, a disk drive, an optical storage device, a solid-state memory, a floppy disk, a flexible disk, a hard disk, a magnetic tape or any other magnetic medium, a compact disk or any other optical medium, a cache memory and/or any other memory chip or module, and/or any other medium from which a computer can read data, instructions and/or code.
- the computing device 600 may also include a random access memory (RAM) 610 and a read-only memory (ROM) 612.
- RAM random access memory
- ROM read-only memory
- the ROM 612 may store programs, utilities, or processes to be executed in a non-volatile manner.
- the RAM 610 may provide volatile data storage and store instructions related to the operation of the computing device 600.
- the computing device 600 may also include a network/bus interface 616 coupled to a data link 618.
- the network/bus interface 616 can be any type of device or system capable of enabling communication with an external device and/or a network, and can include but is not limited to a modem, a network card, an infrared communication device, a wireless communication device and/or a chipset (such as a BluetoothTM device, an IEEE 802.11 device, a WiFi device, a WiMax device, a mobile cellular communication facility, etc.).
- the present disclosure is implemented as a system, apparatus, method, or computer-readable storage medium (e.g., a non-transitory storage medium) as a computer program product.
- a computer program product e.g., a non-transitory storage medium
- the present disclosure may be implemented in various forms, such as a complete hardware embodiment, a complete software embodiment (including firmware, resident software, microprogram code, etc.), or a combination of software and hardware, hereinafter referred to as a "circuit,” "module,” or “system.”
- the present disclosure may also be implemented as a computer program product in any tangible media form having computer-usable program code stored thereon.
- each block in the flowchart or block diagram can represent a module, segment, or portion of program code, which includes one or more executable instructions to implement the specified logical function.
- the functions described by the blocks may not be performed in the order shown in the figures. For example, two blocks connected in the figures can actually be executed simultaneously, or in some cases can be executed in the opposite order of the icons depending on the functions involved.
- each block of the block diagram and/or flowchart, and the combination of blocks in the block diagram and/or flowchart can be implemented by a system based on dedicated hardware, or by a combination of dedicated hardware and computer instructions to perform specific functions or operations.
- An electronic device for intelligent assisted interaction comprising:
- processing circuit being configured to:
- an operating mode of the electronic device Determining an operating mode of the electronic device according to the function setting, and preparing an artificial intelligence (AI) model to be used for the operating mode, the AI model including a speech processing model and/or an image processing model;
- AI artificial intelligence
- the user's reception mode for the inference result is determined, and the inference result is processed according to the determined reception mode.
- an application programming interface (API) related to the inference result is selected to be provided to the user.
- API application programming interface
- the sign language recognition algorithm model is used to infer the collected image information
- the result of the inference is encrypted and provided to the user.
- the frequency domain image information to which the disturbance information and the watermark information are added is subjected to an inverse frequency domain transformation to generate the encrypted image information.
- inference metadata as a result of the inference is provided to the user.
- the image processing model includes a human key point analysis model for extracting key point information of user actions and/or behaviors.
- the result of the inference is encrypted and provided to the user.
- the frequency domain image information to which the disturbance information and the watermark information are added is subjected to an inverse frequency domain transformation to generate the encrypted image information.
- inference metadata as a result of the inference is provided to the user.
- the number of key points extracted using the human body key point analysis model is determined.
- a computer-readable storage medium comprising executable instructions, which, when executed by an information processing device, causes the information processing device to execute the method according to any one of (11) to (19).
- (21) A computer program product comprising a computer program, wherein when the computer program is executed by a processor, the processor is caused to perform the method described in any one of (11) to (19).
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Physics & Mathematics (AREA)
- General Physics & Mathematics (AREA)
- Multimedia (AREA)
- General Engineering & Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Human Computer Interaction (AREA)
- Computer Vision & Pattern Recognition (AREA)
- Evolutionary Computation (AREA)
- Artificial Intelligence (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Data Mining & Analysis (AREA)
- Computing Systems (AREA)
- Software Systems (AREA)
- Medical Informatics (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Acoustics & Sound (AREA)
- Databases & Information Systems (AREA)
- Social Psychology (AREA)
- Psychiatry (AREA)
- Life Sciences & Earth Sciences (AREA)
- Evolutionary Biology (AREA)
- Mathematical Physics (AREA)
- Bioinformatics & Cheminformatics (AREA)
- Bioinformatics & Computational Biology (AREA)
- Computer Hardware Design (AREA)
- Quality & Reliability (AREA)
- User Interface Of Digital Computer (AREA)
Abstract
一种用于智能辅助交互的电子设备,电子设备包括:处理电路,处理电路被配置为:接收用户对电子设备的功能设定(S2010);根据功能设定,确定电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,AI模型包括语音处理模型和/或图像处理模型(S2020);利用所准备的AI模型对采集的语音和/或图像信息进行推理(S2030);根据功能设定,确定用户对推理的结果的接收模式,根据所确定的接收模式对推理的结果进行处理(S2040)。
Description
优先权声明
本申请要求于2024年3月29日递交、申请号为202410374796.0、名称为“用于智能辅助交互的电子设备和方法”的中国专利申请的优先权,其全部内容通过引用并入本文。
本公开涉及计算机领域,更具体地,本公开涉及智能辅助交互。
近年来,伴随着智能化技术的不断发展,尤其是人工智能(AI,Artificial Intelligence)技术的突飞猛进,智能辅助交互设备已经逐渐走进了我们的日常生活。例如,智能语音助手可以接收用户的自然语言输入,通过语音识别、语义分析等基于AI的技术手段,充分理解用户的意图,从而向用户给出对应信息或者提供持续交互。
在下文中给出了关于本公开的简要概述,以便提供关于本公开的一些方面的基本理解。但是,应当理解,这个概述并不是关于本公开的穷举性概述。它并不是意图用来确定本公开的关键性部分或重要部分,也不是意图用来限定本公开的范围。其目的仅仅是以简化的形式给出关于本公开的某些概念,以此作为稍后给出的更详细描述的前序。
本公开的发明人知晓的智能辅助交互系统主要体现为各种语音交互助手的形式,例如智能音箱、聊天机器人等,而缺乏图像类输入的智能化服务。另一方面,图像类输出的交互系统(例如监控系统)的智能化程度较低,例如只能实现检测婴儿啼哭、检测是否有人这样的简单功能,难以完成对复杂场景监控的定制化需求(比如被监测者是否在指定时间学习或运动等需求)。
鉴于以上问题中的一个或多个,本公开提供了一种用于智能辅助交互的电子设备和方法,能够提供语音和图像集成式交互,并且支持用户定制化复杂场景服务设定,提升用户的交互体验。
根据本公开的一个方面,提供了一种用于智能辅助交互的电子设备,所述电子设备包括:处理电路,所述处理电路被配置为:接收用户对所述电子设备的功能设定;根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
根据本公开的另一方面,提供了一种用户端设备,用于与本公开的用于智能辅助交互的电子设备进行交互。该用户端设备包括处理电路,所述处理电路被配置为:显示交互界面,以供用户进行所述电子设备的功能设定;将所述功能设定提供给所述电子设备;以及接收来自所述电子设备的所述推理的结果。
根据本公开的另一方面,提供了一种智能辅助交互方法,包括:接收用户对用于智能辅助交互的电子设备的功能设定;根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
根据本公开的另一方面,提供了一种计算机可读存储介质,包括可执行指令,当所述可执行指令由信息处理装置执行时,使所述信息处理装置执行根据本公开的智能辅助交互方法。
根据本公开的又一方面,提供了一种计算机程序产品,包括计算机程序,所述计算机程序被处理器运行时,使所述处理器执行根据本公开的智能辅助交互方法。
构成说明书的一部分的附图描述了本公开的实施例,并且连同说明书一起用于解释本公开的原理。
参照附图,根据下面的详细描述,可以更清楚地理解本公开,其中:
图1是示出根据本公开的实施例的用于智能辅助交互的电子设备的示例性配置框图以及该电子设备所应用的系统环境;
图2是示出根据本公开的实施例的智能辅助交互方法的示例性流程图;
图3是示出根据本公开的另一实施例的智能辅助交互方法的示例性流程图;
图4是示出根据本公开的实施例的推理元数据的示意图;
图5是示出根据本公开的实施例的图像加密处理的示例性流程图;
图6示出了可以实现根据本发明的实施例的计算设备的示例性配置。
现在将参照附图来详细描述本公开的各种示例性实施例。应注意到:除非另外具体说明,否则在这些实施例中阐述的部件和步骤的相对布置、数字表达式和数值不限制本公开的范围。
同时,应当明白,为了便于描述,附图中所示出的各个部分的尺寸并不是按照实际的比例关系绘制的。
以下对至少一个示例性实施例的描述实际上仅仅是说明性的,决不作为对本公开及其应用或使用的任何限制。
对于相关领域普通技术人员已知的技术、方法和设备可能不作详细讨论,但在适当情况下,所述技术、方法和设备应当被视为说明书的一部分。
在这里示出和讨论的所有示例中,任何具体值应被解释为仅仅是示例性的,而不是作为限制。因此,示例性实施例的其它示例可以具有不同的值。
应注意到:相似的标号和字母在下面的附图中表示类似项,因此,一旦某一项在一个附图中被定义,则在随后的附图中不需要对其进行进一步讨论。
图1示出了根据本公开的实施例的用于智能辅助交互的电子设备1000的示例性配置框图以及该电子设备1000所应用的系统环境。
本公开中所描述的智能辅助交互是指利用人工智能技术等实现用户与智能辅助系统(例如电子设备1000)的交互,从而为用户提供个性化、智能化的服务和帮助的技术。可以支持多种不同的交互方式,包括但不限于语音、图像、手势,可以利用多模态的信息与智能辅助系统进行交互。
如图1所示,电子设备1000与用户端设备1100可通信地耦合,从用户端设备1100接收用于智能辅助交互的功能设定,并根据该功能设定来确定将用于智能辅助交互的电子设备的运行模式,以在该运行模式下执行相应的智能辅助交互动作。另外,电子设备1000可以将执行智能辅助交互动作得到的结果(例如推理结果)发送到用户端设备1100,以使用户获得所需要的信息。
另外,可选地,电子设备1000还可以与云端平台1200可通信地耦合,以从云端平台1200更新用于执行相应智能辅助交互动作的AI模型。另外,在一些情况下,电子设备1000也可以将推理结果等发送给云端平台1200,和/或经由云端平台1200接收功能设定。
另外,可选地,云端平台1200与用户端设备1100也可以可通信地耦合。用户端设备1100可以将用于智能辅助交互的功能设定发送到云端平台1200,并经由云端平台1200接收来自电子设备1000的智能辅助交互动作的推理结果。
如图1所示,在一些实施例中,电子设备1000可以包括处理电路1010。电子设备1000的处理电路1010提供电子设备1000的各种功能。在一些实施例中,电子设备1000的处理电路1010可以被配置为执行用于电子设备1000的方法。另外,用户端设备1100也可以包括处理电路(未图示),用于提供用户端设备1100的各种功能。
处理电路1010(或用户端设备1100的处理电路)可以指在计算系统中执行功能的数字电路系统、模拟电路系统或混合信号(模拟和数字的组合)电路系统的各种实现。处理电路可以包括例如诸如集成电路(IC)、专用集成电路(ASIC)这样的电路、单独处理器核心的部分或电路、整个处理器核心、单独的处理器、诸如现场可编程门阵列(FPGA)的可编程硬件设备、和/或包括多个处理器的系统。
在一些实施例中,处理电路1010可以包括信息接收单元1020、推理单元1030、模型准备单元1040和后处理单元,配置为执行后述图2中所示的方法2000中的相应步骤。
在一些实施例中,电子设备1000还可以包括存储器(未图示)。电子设备1000的存储器可以存储由处理电路1010产生的信息以及用于电子设备1000操作的程序和数据。存储器可以是易失性存储器和/或非易失性存储器。例如,存储器可以包括但不限于随机存取存储器(RAM)、动态随机存取存储器(DRAM)、静态随机存取存储器(SRAM)、只读存储器(ROM)以及闪存存储器。
另外,电子设备1000可以以芯片级来实现,或者也可以通过包括其它外部部件而以设备级来实现。在一些实施例中,电子设备1000可以作为整机实现,并且还可以包括多根天线。
应当理解,上述各个单元仅是根据其所实现的具体功能所划分的逻辑模块,而不是用于限制具体的实现方式。在实际实现时,上述各个单元可被实现为独立的物理实体,或者也可由单个实体(例如,处理器(CPU或DSP等)、集成电路等)来实现。
图2示出了根据本公开的实施例的智能辅助交互方法2000的示例性流程图,该方法2000例如可以由图1所示的电子设备1000来实现。
如图2所示,在S2010中,信息接收单元1020接收用户对电子设备1000的功能设定。该功能设定例如用于设定用户期望的电子设备1000的运行模式、推理结果的接收方式等,使电子设备1000在该运行模式下运行以实现用户期望的辅助交互功能,并以用户期望的接收方式返回推理结果。
在一些实施例中,用户可以在用户端设备1100处(例如通过自己的本地终端设备(例如电脑、平板电脑、智能手机、智能手表等))下载本地客户端程序,通过用户端设备1100进行功能设定。
在一些实施例中,用户端设备1100的处理电路可以被配置为显示交互界面,以供用户进行电子设备1000的功能设定。例如,交互界面中可以显示可选的运行模式的菜单选项,以供用户选择和设定电子设备1000的期望的运行模式。
在一些实施例中,运行模式可以包括用户定制化服务模式和开发者调试模式。在用户定制化服务模式下,可以提供用户所需的典型服务,在开发者调试模式下,可以提供开发者开发或调试所需的辅助工具/信息。例如,可以在交互界面中显示用户定制化服务模式和开发者调制模式的菜单选项,以供用户根据需要进行选择。
在一些实施例中,响应于用户选择用户定制化服务模式,可以进一步在交互界面中显示二级菜单,供用户选择用户定制化服务模式下的子模式,例如包括游戏模式、居家模式、智能助手模式等。接下来,响应于用于在二级菜单中选择相应的子模式,可以进一步在交互界面中显示三级菜单,以供用户选择该子模式下的进一步细分的服务模式。例如,在游戏模式下,可以包括瑜伽练习(监测瑜伽动作是否正确)、唱跳练习(监测唱跳练习是否符合标准并辅助进行改进)等模式,在居家模式下,可以包括监测学生活动(监测在所设定的学习时段是否学习)、健康监测(通过捕捉动作或身形来评估身体健康状况)等模式;在智能助手模式下,可以包括残障人士辅助模式,例如可以针对残障人士(尤其是聋哑人士)进行手语、动作识别等(例如通过手语控制音乐播放)。
另外,在一些实施例中,响应于用户选择开发者调试模式,可以在交互界面的二级菜单中显示该模式下的多种子模式。例如,按开发者类型分类,子模式可以包括游戏开发者模式和算法开发者模式等。针对游戏开发者模式,可以向用户提供人体位置信息、人体动作识别结果等中间结果,用于辅助用户进行游戏开发;针对算法开发者模式,可以向用户提供人体关键点信息(例如二维或三维关键点信息)等,用于辅助相应的算法开发。用户可以在该二级菜单中选择所需的子模式。另外,也可以按场景类型分类,子模式包括分类模式、检测模式、文本识别模式、图像分割模式、语义分割模式等,供开发者根据需要进行选择。此外,也可以针对子模式进行进一步的分类,供用户进行更细化的选择。例如,检测模式可以进一步分类为常规检测模式、关键点检测模式、人脸检测模式等。
应当理解,本公开中的运行模式不限于用户定制化服务模式和开发者调试模式,也可以根据需要设计其他运行模式。另外,用户定制化服务模式和开发者调试模式下的各子模式也仅为示例,本领域技术人员根据实际需求,可以基于其他分类方式来确定子模式及其进一步细分的模式,以满足用户和开发者的不同需求。
根据本公开的实施例,用户能够通过用户端设备1100来进行功能设定,以设定电子设备1000的运行模式。在一些实施例中,用户端设备1100的处理电路可以被配置为将功能设定提供给电子设备1000,由电子设备1000的信息接收单元1020接收该功能设定,并据此确定电子设备1000的运行模式以进行后续处理。
如图2所示,在S2020中,模型准备单元1030根据在S2010中接收的功能设定,确定电子设备1000的运行模式,并准备要用于该运行模式的AI模型。该AI模型包括语音处理模型和/或图像处理模型。
在一些实施例中,电子设备1000中的存储器中可以预先存储多个AI模型,包括多种类型的语音处理模型和图像处理模型。根据用户设定的运行模式,模型准备单元1030可以从预先存储的多个AI模型中选择该运行模式所需的AI模型。
例如,在用户设定的运行模式为用户定制化服务模式下的瑜伽练习模式的情况下,可以从多个AI模型中选择相应的图像处理模型(例如人体关键点分析模型等),用于分析用户进行瑜伽练习时的动作,并进行相应的推理。作为另一个例子,在所设定的运行模式为用户定制化服务模式下的唱跳练习模式的情况下,可以从多个AI模型中选择相应的语音处理模型和图像处理模型,分别对用户的声音和动作进行分析和推理。另外,在所设定的运行模式为开发者调试模式的情况下,也可以从多个AI模型中选择开发者所需的AI模型。
另外,在一些实施例中,预先准备的多个AI模型也可以存储(或部分存储)在云端平台1200。根据用户所设定的运行模式,利用云端平台1200可以进行AI模型的更新。例如,在根据用户所设定的运行模式确定的AI模型没有存储在电子设备1000中的情况下,可以从云端平台1200将相应的AI模型下载到电子设备1000中,以供后续使用。另外,在完成了AI模型的更新后,电子设备1000可以离线使用而不依赖于网络条件,从而能够扩大使用范围。
接下来,在S2030中,推理单元1040利用所准备的AI模型对采集的语音和/或图像信息进行推理。
在一些实施例中,电子设备1000可以包括语音/图像采集单元,用于采集待推理的语音和/或图像信息,例如用户的瑜伽动作、唱跳练习的语音和动作等。在另一些实施例中,电子设备1000也可以与语音/图像采集设备(例如麦克风、摄像头等)可通信地连接,接收语音/图像采集单元采集的语音和/或图像信息。
例如,在所设定的运行模式为用户定制化服务模式下的瑜伽练习模式的情况下,可以采集用户的瑜伽动作图像/视频,利用所准备的人体关键点分析模型来提取关键点,从而对用户的瑜伽动作进行推理,以检测瑜伽动作是否正确。另外,在所设定的运行模式为用户定制化服务模式下的唱跳练习场景的情况下,可以采集用户进行唱跳练习时的语音和图像/视频,并利用相应的语音处理类AI模型推理歌唱部分的练习情况,利用图像处理类AI模型推理跳舞部分的练习情况,从而辅助用户的唱跳练习。
接下来,在S2040中,后处理单元1050根据用户的功能设定,确定用户对推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。在本公开中,将对推理的结果进行的处理都称为“后处理”。
在一些实施例中,可以在用户端设备1100显示的交互界面中,通过功能设定来设定对推理的结果的期望接收模式(例如通过菜单选项),包括语音接收模式、图像接收模式、元数据接收模式中的一种或多种。该接收模式可以包括在功能设定中一并提供给电子设备1000的信息接收单元1020。
在另一些实施例中,也可以根据用户的功能设定中的运行模式来确定接收模式。例如,响应于运行模式为用户定制化服务模式,可以将推理的结果转化为可读信息以提供给所述用户。可读信息例如可以是图像信息、语音信息、文字信息等。由此,用户可以获得与所设定的运行模式对应的处理结果。另外,响应于运行模式为开发者调试模式,可以选择与推理的结果有关的应用程序接口API以提供给所述用户,以向用户(开发者)提供更底层的API服务,方便开发者进行推理元数据查询(如通过获取坐标点API获取二维、三维关键点坐标,或通过获取识别结果API获取被测者目前的动作识别结果等),加速其开发流程。
根据本公开的用于智能辅助交互的电子设备和方法,能够支持语音和图像集成式交互,能够提升用户的智能交互产品的使用体验。另外,可以提供用户定制化复杂场景服务设定,根据客户的需求提供相应的交互体验,提供多样化的功能支持。另外,在开发者调试模式下,还支持二次开发,向开发者提供所需的API服务,能够加速复杂应用的开发进程。
接下来,参照图3,描述根据本公开的另一实施例的智能辅助交互方法3000的示例性流程图,该方法3000例如也可以由图1所示的电子设备1000来实现。
如图3所示,在步骤S3010中,判断是否由用户设定了运行模式。该步骤例如可以通过是否接收到来自用户端设备1100或云端平台1200的功能设定,或者所接收到的功能设定中是否包括关于运行模式的设定来判断。
在一些实施例中,在步骤S3010中判断为“否”的情况下,例如接收到的功能设定中没有关于运行模式的设定,电子设备1000的运行模式可以为“残障人士助理模式”,该模式例如可以是电子设备1000的默认的运行模式。
另外,在步骤S3010中判断为“是”的情况下,电子设备1000根据用户设定的运行模式,在S3015中进一步判断是否为“用户定制化服务模式”(简称为“用户模式”)。在S3015中判断为“是”的情况下,电子设备1000的运行模式被设定为用户模式。另外,在S3015中判断为“否”的情况下,电子设备的运行模式被设定为“开发者调试模式”(简称为“开发者模式”)。
接下来,在根据来自用户的运行模式的设定而完成了电子设备1000的运行模式的设定后,方法3000进入到步骤S3020,根据所设定的运行模式来准备相应的AI模型。该步骤与根据图2描述的步骤S2020对应。
在一些实施例中,可以预先准备手语识别算法模型,用于识别采集的残障人士的图像/视频信息中的手语,以供在运行模式为默认的残障人士助理模式的情况下使用。另外,手语识别算法模型例如可以预先存储在电子设备1000的存储器中,而无需从云端平台1300下载更新,从而支持离线使用残障人士助理模式。由此,能够实现残障人士(尤其是聋哑人士)与电子设备的智能辅助交互,提升使用体验。此外,也可以预先准备语音识别算法模型,用于识别例如盲人等残障人士的语音交互。
本公开中作为示例,将“残障人士助理模式”设置为默认的运行模式,残障人士(如聋哑人士、盲人等)无需另外设定运行模式就能直接与电子设备1000进行交互,从而使得残障人士使用电子设备1000时的操作更便利。另外,应当理解,本领域技术人员也可以根据实际需求,将其他运行模式设置为默认模式,或者也可以不设置默认的运行模式。另外,运行模式的种类也不限于以上列举的三种,可以包括其他运行模式,或者对运行模式进行其他方式的分类。
接下来,在步骤S3025中,利用所准备的AI模型对采集的语音和/或图像信息进行推理。该步骤与根据图2描述的步骤S2030对应。
例如,可以根据所设定的运行模式,在预定时间段内或按照预定的周期,对监测对象的语音、动作、手势等进行采集,以用于进行推理。
例如,在运行模式为监测学生活动模式的情况下,用户(例如家长)可以通过用户端设备1100进一步设定监测时间段,作为针对电子设备1000的功能设定,并提供给电子设备1000。电子设备1000根据所设定的监测时间段控制其图像采集单元对检测对象(例如学生)进行图像采集,或者控制与其连接的图像采集设备进行图像采集,从而监测学生在所设定的学习时段是否学习。
另外,取决于不同的运行模式,监测对象可以与用户相同,也可以不同。在上述监测学生活动模式的示例中,检测对象与用户是不同的。作为另一示例,在运行模式为唱跳练习模式的情况下,用户和监测对象也可以相同,电子设备1000获取所设定的时间段内的用户的唱跳练习视频,以监测唱跳练习是否符合标准并辅助进行改进。
接下来,在步骤S3030中,判断运行模式是否为开发者模式。在一些实施例中,该步骤也可以与步骤S3015的判断处理合并,或者利用步骤S3015中的判断结果。在步骤S3030中判断为“是”的情况下,进入到步骤S3040以向用户提供API服务。另外,在步骤S3030中判断为“否”的情况下,将步骤S3025中推理得到的推理的结果转化为可读信息提供给用户。
在一些实施例中,在步骤S3035中,可以进一步确定用户对于推理的结果的接收模式。在一些实施例中,接收模式可以根据用户设定的运行模式来确定。在另一些实施例中,也可以由用户在用户端设备1100处设定接收模式。在步骤S3035中判断接收模式是否为图像接收模式和语音接收模式中的至少一种。在步骤S3035中判断为“否”的情况下,确定不是图像接收模式,也不是语音接收模式,则将作为推理的结果的推理元数据(无图模式)提供给用户。
图4示出了根据本公开的实施例的推理元数据的示意图。该图是利用人体关键点分析模型对人体图像/视频提取的人体关键点的信息。
如图4所示,该推理元数据示出了16个人体关键点,以供用户(开发者)利用该人体关键点信息进行后续的开发等处理。由于无法从推理元数据反推出用户的图像,因此可以保证用户的隐私安全。
在一些实施例中,可以根据所设定的运行模式,确定利用人体关键点分析模型提取的关键点的数目。作为示例,关键点的数目可以设定为16至33个。在运行模式为针对简单场景的监控模式的情况下,可以设定较少的关键点;在运行模式为针对复杂场景的监控模式的情况下,例如,监控被监测者是否在指定时间内学习或运动等,则可以选取较多的关键点,使得推理的结果更准确。
返回参考图3,在步骤S3035中判断为“是”的情况下,考虑到图像/语音信息可能涉及用户的隐私,因此可以在步骤S3050中对推理的结果进行加密处理以提供给用户。由此,即使攻击者截取到推理的结果,由于该结果进行了加密处理,也无法盗用或篡改用户的图像/语音信息。另外,在接收模式不是语音/图像接收模式的情况下,由于仅向用户提供推理元数据,而根据元数据无法反推出用户图像/语音,因此也可以保证用户的隐私安全。
接下来,参照图5,描述根据本公开的实施例的图像加密处理5000的示例性流程图。该图像加密处理5000例如可以对应于图3所示的步骤S3050中提供加密图像的处理。
在步骤S5010中,对推理的结果(例如热力图(heatmap))进行解码以获得频域图像。该解码例如是利用AI模型对图像信息进行推理的逆过程。另外,频域图像是对时域图像进行频域变换得到的,频域变换可以是傅里叶变换、离散余弦变换(DCT)等频域变换。
接下来,在步骤S5020中,对解码所得到的频域图像附加扰动信息。在一些实施例中,例如通过如下方式来生成扰动。
对数据x和一个错误标签l,寻找一个最小的扰动r*,使得分类器f将x+r*错误分类为l,即
s.t.f(x+r)=l,x+r∈[0,1]m
其中,x+r∈[0,1]m表示m维的x+r的归一化后的取值在m维的[0,1]的范围内。由此,生成的扰动r*是对抗扰动,扰动后的数据x+r*是对抗样本。
在一些实施例中,上述问题可以转化为求解如下优化问题:
s.t.f(x+r)=l,x+r∈[0,1]m
其中L是损失函数,θ是模型的参数,c是经验常数。在训练过程中,使得损失函数L最小,就得到了指定错误标签l下的噪声分布数据r。由此,将生成的扰动添加频域图像,以得到加扰后的频域图像。
接下来,在步骤S5030中,对步骤S5020中得到的加扰后的频域图像附加水印。在该步骤中,为了简化说明,如下公式中省略了步骤S5020中附加的扰动。假设将时域图像和时域水印分别表示为f1和f2,经频域变换后的频域图像为F1。F1例如是对原图进行二维离散傅里叶变换(如下公式中表示为“FFT2”)得到的,则
F1=FFT2(f1)
F1=FFT2(f1)
为了让水印的信息能在频域上尽量均匀分布,引入随机变量r(f)来对时域水印进行变换,即,变换后的时域水印可以表示为
f′2=r(f2)
f′2=r(f2)
接下来,引入能量系数α对水印和频域图像进行合成,可以得到附加水印后的频域图像
F=F1+αf′2
F=F1+αf′2
接下来,在步骤S5040中,对上述得到的F进行逆频域变换(例如逆二维傅里叶变换,表示为“IFFT2”),可以得到附加了水印的时域图像f,即步骤S5050中生成的加密图像。
f=IFFT2(F)
f=IFFT2(F)
根据图像加密处理5000生成的加密图像由于经过了附加扰动和附加水印的处理,因此能够确保用户的隐私安全,即使信息被窃取,也无法被其他AI算法利用(因为附加了对抗扰动),也能对信息进行追踪溯源(因为附加了水印)。
另外,如果希望解除水印,则首先对f频域变换得到F=FFT2(f),然后减去原图的频域表示F1并且做随机变换r的逆变换r-1,可以还原水印的时域表示f2,以用于水印解除。
f2=r-1(F-F1)/α
f2=r-1(F-F1)/α
另外,以上例示了图像加密处理,应当理解,对于作为推理的结果而提供给用户的语音信息,也可以进行语音加密处理以保护用户的隐私安全。语音加密的方法可以采用任何合适的方法,本公开对此没有限定。
图6示出了能够实现根据本发明的实施例的计算设备600的示例性配置。
计算设备600是能够应用本发明的上述方面的硬件设备的实例。计算设备600可以是被配置为执行处理和/或计算的任何机器。计算设备600可以是但不限制于工作站、服务器、台式计算机、膝上型计算机、平板计算机、个人数据助手(PDA)、智能电话、车载计算机或以上组合。
如图6所示,计算设备600可以包括可以经由一个或多个接口与总线602连接或通信的一个或多个元件。总线602可以包括但不限于,工业标准架构(Industry Standard Architecture,ISA)总线、微通道架构(Micro Channel Architecture,MCA)总线、增强ISA(EISA)总线、视频电子标准协会(VESA)局部总线、以及外设组件互连(PCI)总线等。计算设备600可以包括例如一个或多个处理器604、一个或多个输入设备606以及一个或多个输出设备608。一个或多个处理器604可以是任何种类的处理器,并且可以包括但不限于一个或多个通用处理器或专用处理器(诸如专用处理芯片)。处理器602例如可以对应于图1中的处理器1010,被配置为实现本公开的用于智能辅助交互的电子设备1000的各单元的功能。另外,处理器602也可以对应于图1中的用户端设备1100的处理电路,被配置为与电子设备1000进行交互。输入设备506可以是能够向计算设备输入信息的任何类型的输入设备,并且可以包括但不限于鼠标、键盘、触摸屏、麦克风和/或远程控制器。输出设备508可以是能够呈现信息的任何类型的设备,并且可以包括但不限于显示器、扬声器、视频/音频输出终端、振动器和/或打印机。
计算设备600还可以包括或被连接至非暂态存储设备614,该非暂态存储设备614可以是任何非暂态的并且可以实现数据存储的存储设备,并且可以包括但不限于盘驱动器、光存储设备、固态存储器、软盘、柔性盘、硬盘、磁带或任何其他磁性介质、压缩盘或任何其他光学介质、缓存存储器和/或任何其他存储芯片或模块、和/或计算机可以从其中读取数据、指令和/或代码的其他任何介质。计算设备600还可以包括随机存取存储器(RAM)610和只读存储器(ROM)612。ROM 612可以以非易失性方式存储待执行的程序、实用程序或进程。RAM 610可提供易失性数据存储,并存储与计算设备600的操作相关的指令。计算设备600还可包括耦接至数据链路618的网络/总线接口616。网络/总线接口616可以是能够启用与外部装置和/或网络通信的任何种类的设备或系统,并且可以包括但不限于调制解调器、网络卡、红外线通信设备、无线通信设备和/或芯片集(诸如蓝牙TM设备、IEEE 802.11设备、WiFi设备、WiMax设备、移动蜂窝通信设施等)。
应当理解,本说明书中“实施例”或类似表达方式的引用是指结合该实施例所述的特定特征、结构、或特性系包括在本公开的至少一具体实施例中。因此,在本说明书中,“在本公开的实施例中”及类似表达方式的用语的出现未必指相同的实施例。
本领域技术人员应当知道,本公开被实施为一系统、装置、方法或作为计算机程序产品的计算机可读存储介质(例如非瞬态存储介质)。因此,本公开可以实施为各种形式,例如完全的硬件实施例、完全的软件实施例(包括固件、常驻软件、微程序代码等),或者也可实施为软件与硬件的实施形式,在以下会被称为“电路”、“模块”或“系统”。此外,本公开也可以任何有形的媒体形式实施为计算机程序产品,其具有计算机可使用程序代码存储于其上。
本公开的相关叙述参照根据本公开具体实施例的系统、装置、方法及计算机程序产品的流程图和/或框图来进行说明。可以理解每一个流程图和/或框图中的每一个块,以及流程图和/或框图中的块的任何组合,可以使用计算机程序指令来实施。这些计算机程序指令可供通用型计算机或特殊计算机的处理器或其它可编程数据处理装置所组成的机器来执行,而指令经由计算机或其它可编程数据处理装置处理以便实施流程图和/或框图中所说明的功能或操作。
在附图中显示根据本公开各种实施例的系统、装置、方法及计算机程序产品可实施的架构、功能及操作的流程图及框图。因此,流程图或框图中的每个块可表示一模块、区段、或部分的程序代码,其包括一个或多个可执行指令,以实施指定的逻辑功能。另外应当注意,在某些其它的实施例中,块所述的功能可以不按图中所示的顺序进行。举例来说,两个图示相连接的块事实上也可以同时执行,或根据所涉及的功能在某些情况下也可以按图标相反的顺序执行。此外还需注意,每个框图和/或流程图的块,以及框图和/或流程图中块的组合,可藉由基于专用硬件的系统来实施,或者藉由专用硬件与计算机指令的组合,来执行特定的功能或操作。
以上已经描述了本公开的各实施例,上述说明是示例性的,并非穷尽性的,并且也不限于所披露的各实施例。在不偏离所说明的各实施例的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实施例的原理、实际应用或对市场技术的技术改进,或者使本技术领域的其它普通技术人员能理解本文披露的各实施例。
注意,本说明书中公开的技术可以具有以下配置。
(1)一种用于智能辅助交互的电子设备,所述电子设备包括:
处理电路,所述处理电路被配置为:
接收用户对所述电子设备的功能设定;
根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;
利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及
根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
(2)根据(1)所述的电子设备,其中,所述运行模式包括用户定制化服务模式和开发者调试模式,所述处理电路被进一步配置为:
响应于所述运行模式确定为所述用户定制化服务模式,将所述推理的结果转化为可读信息以提供给所述用户;以及
响应于所述运行模式确定为所述开发者调试模式,选择与所述推理的结果有关的应用程序接口API以提供给所述用户。
(3)根据(1)所述的电子设备,其中,默认的运行模式为残障人士助理模式,所述处理电路被进一步配置为:
在所述残障人士助理模式下,利用手语识别算法模型对采集的图像信息进行推理;以及
将所述推理的结果转化为可读信息以提供给所述用户。
(4)根据(1)所述的电子设备,其中,所述处理电路被进一步配置为:
根据所述功能设定,确定用户对所述推理的结果的接收模式是否为图像接收模式和语音接收模式中的至少一种;以及
响应于所确定的接收模式为图像接收模式和语音接收模式中的至少一种,对所述推理的结果进行加密处理以提供给所述用户。
(5)根据(4)所述的电子设备,其中,在所述接收模式为图像接收模式的情况下,所述加密处理包括:
对根据所述推理的结果生成的频域图像信息附加扰动信息和水印信息;以及
对附加了扰动信息和水印信息的频域图像信息进行逆频域变换,以生成所述加密的图像信息。
(6)根据(4)所述的电子设备,其中,所述处理电路被进一步配置为:
响应于所确定的接收模式不是图像接收模式和语音接收模式,将作为所述推理的结果的推理元数据提供给所述用户。
(7)根据(1)所述的电子设备,其中,所述处理电路被进一步配置为:
利用云端平台更新所述AI模型;以及
利用更新后的AI模型,离线地对采集的语音和/或图像信息进行推理。
(8)根据(1)所述的电子设备,其中,所述图像处理模型包括人体关键点分析模型,用于提取用户动作和/或行为的关键点信息。
(9)根据(8)所述的电子设备,其中,所述处理电路被进一步配置为:
根据所述运行模式,确定利用所述人体关键点分析模型提取的关键点的数目。
(10)一种用户端设备,用于与根据上述(1)至(9)中任一项所述的电子设备进行交互,所述用户端设备包括处理电路,所述处理电路被配置为:
显示交互界面,以供用户进行所述电子设备的功能设定;
将所述功能设定提供给所述电子设备;以及
接收来自所述电子设备的所述推理的结果。
(11)一种智能辅助交互方法,包括:
接收用户对用于智能辅助交互的电子设备的功能设定;
根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;
利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及
根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
(12)根据(11)所述的方法,其中,所述运行模式包括用户定制化服务模式和开发者调试模式,所述方法还包括:
响应于所述运行模式确定为所述用户定制化服务模式,将所述推理的结果转化为可读信息以提供给所述用户;以及
响应于所述运行模式确定为所述开发者调试模式,选择与所述推理的结果有关的应用程序接口API以提供给所述用户。
(13)根据(11)所述的方法,其中,默认的运行模式为残障人士助理模式,所述方法还包括:
在所述残障人士助理模式下,利用手语识别算法模型对采集的图像信息进行推理;以及
将所述推理的结果转化为可读信息以提供给所述用户。
(14)根据(11)所述的方法,其中,所述方法还包括:
根据所述功能设定,确定用户对所述推理的结果的接收模式是否为图像接收模式和语音接收模式中的至少一种;以及
响应于所确定的接收模式为图像接收模式和语音接收模式中的至少一种,对所述推理的结果进行加密处理以提供给所述用户。
(15)根据(14)所述的方法,其中,在所述接收模式为图像接收模式的情况下,所述加密处理包括:
对根据所述推理的结果生成的频域图像信息附加扰动信息和水印信息;以及
对附加了扰动信息和水印信息的频域图像信息进行逆频域变换,以生成所述加密的图像信息。
(16)根据(14)所述的方法,其中,所述方法还包括:
响应于所确定的接收模式不是图像接收模式和语音接收模式,将作为所述推理的结果的推理元数据提供给所述用户。
(17)根据(11)所述的方法,还包括:
利用云端平台更新所述AI模型;以及
利用更新后的AI模型,离线地对采集的语音和/或图像信息进行推理。
(18)根据(11)所述的方法,其中,所述图像处理模型包括人体关键点分析模型,用于提取用户动作和/或行为的关键点信息。
(19)根据(18)所述的方法,还包括:
根据所述运行模式,确定利用所述人体关键点分析模型提取的关键点的数目。
(20)一种计算机可读存储介质,包括可执行指令,当所述可执行指令由信息处理装置执行时,使所述信息处理装置执行根据(11)至(19)中任一项所述的方法。
(21)一种计算机程序产品,包括计算机程序,所述计算机程序被处理器运行时,使所述处理器执行(11)至(19)中的任一项所述的方法。
Claims (21)
- 一种用于智能辅助交互的电子设备,所述电子设备包括:处理电路,所述处理电路被配置为:接收用户对所述电子设备的功能设定;根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
- 根据权利要求1所述的电子设备,其中,所述运行模式包括用户定制化服务模式和开发者调试模式,所述处理电路被进一步配置为:响应于所述运行模式确定为所述用户定制化服务模式,将所述推理的结果转化为可读信息以提供给所述用户;以及响应于所述运行模式确定为所述开发者调试模式,选择与所述推理的结果有关的应用程序接口API以提供给所述用户。
- 根据权利要求1所述的电子设备,其中,默认的运行模式为残障人士助理模式,所述处理电路被进一步配置为:在所述残障人士助理模式下,利用手语识别算法模型对采集的图像信息进行推理;以及将所述推理的结果转化为可读信息以提供给所述用户。
- 根据权利要求1所述的电子设备,其中,所述处理电路被进一步配置为:根据所述功能设定,确定用户对所述推理的结果的接收模式是否为图像接收模式和语音接收模式中的至少一种;以及响应于所确定的接收模式为图像接收模式和语音接收模式中的至少一种,对所述推理的结果进行加密处理以提供给所述用户。
- 根据权利要求4所述的电子设备,其中,在所述接收模式为图像接收模式的情况下,所述加密处理包括:对根据所述推理的结果生成的频域图像信息附加扰动信息和水印信息;以及对附加了扰动信息和水印信息的频域图像信息进行逆频域变换,以生成所述加密的图像信息。
- 根据权利要求4所述的电子设备,其中,所述处理电路被进一步配置为:响应于所确定的接收模式不是图像接收模式和语音接收模式,将作为所述推理的结果的推理元数据提供给所述用户。
- 根据权利要求1所述的电子设备,其中,所述处理电路被进一步配置为:利用云端平台更新所述AI模型;以及利用更新后的AI模型,离线地对采集的语音和/或图像信息进行推理。
- 根据权利要求1所述的电子设备,其中,所述图像处理模型包括人体关键点分析模型,用于提取用户动作和/或行为的关键点信息。
- 根据权利要求8所述的电子设备,其中,所述处理电路被进一步配置为:根据所述运行模式,确定利用所述人体关键点分析模型提取的关键点的数目。
- 一种用户端设备,用于与根据权利要求1至9中任一项所述的电子设备进行交互,所述用户端设备包括处理电路,所述处理电路被配置为:显示交互界面,以供用户进行所述电子设备的功能设定;将所述功能设定提供给所述电子设备;以及接收来自所述电子设备的所述推理的结果。
- 一种智能辅助交互方法,包括:接收用户对用于智能辅助交互的电子设备的功能设定;根据所述功能设定,确定所述电子设备的运行模式,并准备要用于该运行模式的人工智能AI模型,所述AI模型包括语音处理模型和/或图像处理模型;利用所准备的所述AI模型对采集的语音和/或图像信息进行推理;以及根据所述功能设定,确定所述用户对所述推理的结果的接收模式,根据所确定的接收模式对所述推理的结果进行处理。
- 根据权利要求11所述的方法,其中,所述运行模式包括用户定制化服务模式和开发者调试模式,所述方法还包括:响应于所述运行模式确定为所述用户定制化服务模式,将所述推理的结果转化为可读信息以提供给所述用户;以及响应于所述运行模式确定为所述开发者调试模式,选择与所述推理的结果有关的应用程序接口API以提供给所述用户。
- 根据权利要求11所述的方法,其中,默认的运行模式为残障人士助理模式,所述方法还包括:在所述残障人士助理模式下,利用手语识别算法模型对采集的图像信息进行推理;以及将所述推理的结果转化为可读信息以提供给所述用户。
- 根据权利要求11所述的方法,其中,所述方法还包括:根据所述功能设定,确定用户对所述推理的结果的接收模式是否为图像接收模式和语音接收模式中的至少一种;以及响应于所确定的接收模式为图像接收模式和语音接收模式中的至少一种,对所述推理的结果进行加密处理以提供给所述用户。
- 根据权利要求14所述的方法,其中,在所述接收模式为图像接收模式的情况下,所述加密处理包括:对根据所述推理的结果生成的频域图像信息附加扰动信息和水印信息;以及对附加了扰动信息和水印信息的频域图像信息进行逆频域变换,以生成所述加密的图像信息。
- 根据权利要求14所述的方法,其中,所述方法还包括:响应于所确定的接收模式不是图像接收模式和语音接收模式,将作为所述推理的结果的推理元数据提供给所述用户。
- 根据权利要求11所述的方法,还包括:利用云端平台更新所述AI模型;以及利用更新后的AI模型,离线地对采集的语音和/或图像信息进行推理。
- 根据权利要求11所述的方法,其中,所述图像处理模型包括人体关键点分析模型,用于提取用户动作和/或行为的关键点信息。
- 根据权利要求18所述的方法,还包括:根据所述运行模式,确定利用所述人体关键点分析模型提取的关键点的数目。
- 一种计算机可读存储介质,包括可执行指令,当所述可执行指令由信息处理装置执行时,使所述信息处理装置执行根据权利要求11至19中任一项所述的方法。
- 一种计算机程序产品,包括计算机程序,所述计算机程序被处理器运行时,使所述处理器执行根据权利要求11至19中的任一项所述的方法。
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN202410374796.0 | 2024-03-29 | ||
| CN202410374796.0A CN120727003A (zh) | 2024-03-29 | 2024-03-29 | 用于智能辅助交互的电子设备和方法 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2025201223A1 true WO2025201223A1 (zh) | 2025-10-02 |
Family
ID=97165535
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2025/084290 Pending WO2025201223A1 (zh) | 2024-03-29 | 2025-03-24 | 用于智能辅助交互的电子设备和方法 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN120727003A (zh) |
| WO (1) | WO2025201223A1 (zh) |
Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160179655A1 (en) * | 2014-12-18 | 2016-06-23 | Red Hat, Inc. | Automatic Switch To Debugging Mode |
| CN111552630A (zh) * | 2020-03-24 | 2020-08-18 | 北京声智科技有限公司 | 技能调试的方法、装置及存储介质 |
| CN113760558A (zh) * | 2020-06-01 | 2021-12-07 | 阿里巴巴集团控股有限公司 | 数据处理方法、设备及计算机存储介质 |
| CN113868401A (zh) * | 2021-10-18 | 2021-12-31 | 深圳追一科技有限公司 | 数字人的交互方法、装置、电子设备及计算机存储介质 |
| CN117122887A (zh) * | 2023-08-28 | 2023-11-28 | 段芸莱 | 一种ai教练系统 |
| CN117383378A (zh) * | 2023-09-27 | 2024-01-12 | 电子科技大学长三角研究院(湖州) | 安全控制系统及控制调整方法 |
| CN117762257A (zh) * | 2023-12-28 | 2024-03-26 | 华中科技大学 | 一种基于多模态大语言模型的人机交互系统及方法 |
-
2024
- 2024-03-29 CN CN202410374796.0A patent/CN120727003A/zh active Pending
-
2025
- 2025-03-24 WO PCT/CN2025/084290 patent/WO2025201223A1/zh active Pending
Patent Citations (7)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20160179655A1 (en) * | 2014-12-18 | 2016-06-23 | Red Hat, Inc. | Automatic Switch To Debugging Mode |
| CN111552630A (zh) * | 2020-03-24 | 2020-08-18 | 北京声智科技有限公司 | 技能调试的方法、装置及存储介质 |
| CN113760558A (zh) * | 2020-06-01 | 2021-12-07 | 阿里巴巴集团控股有限公司 | 数据处理方法、设备及计算机存储介质 |
| CN113868401A (zh) * | 2021-10-18 | 2021-12-31 | 深圳追一科技有限公司 | 数字人的交互方法、装置、电子设备及计算机存储介质 |
| CN117122887A (zh) * | 2023-08-28 | 2023-11-28 | 段芸莱 | 一种ai教练系统 |
| CN117383378A (zh) * | 2023-09-27 | 2024-01-12 | 电子科技大学长三角研究院(湖州) | 安全控制系统及控制调整方法 |
| CN117762257A (zh) * | 2023-12-28 | 2024-03-26 | 华中科技大学 | 一种基于多模态大语言模型的人机交互系统及方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN120727003A (zh) | 2025-09-30 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| Cruciani et al. | Feature learning for human activity recognition using convolutional neural networks: A case study for inertial measurement unit and audio data | |
| CN110544488B (zh) | 一种多人语音的分离方法和装置 | |
| CN107369196B (zh) | 表情包制作方法、装置、存储介质及电子设备 | |
| CN103140862B (zh) | 用户界面系统及其操作方法 | |
| US10242253B2 (en) | Detection apparatus, detection method, and computer program product | |
| CN104361896B (zh) | 语音质量评价设备、方法和系统 | |
| CN110363084A (zh) | 一种上课状态检测方法、装置、存储介质及电子 | |
| CN110741387B (zh) | 人脸识别方法、装置、存储介质及电子设备 | |
| CN110659412A (zh) | 用于在电子设备中提供个性化服务的方法和设备 | |
| CN115620384A (zh) | 模型训练方法、眼底图像预测方法及装置 | |
| WO2020238321A1 (zh) | 用于识别年龄的方法和装置 | |
| CN108510084A (zh) | 用于生成信息的方法和装置 | |
| CN114724237A (zh) | 动作识别方法、装置、存储介质及电子设备 | |
| TW202113685A (zh) | 人臉辨識的方法及裝置 | |
| CN116596748A (zh) | 图像风格化处理方法、装置、设备、存储介质和程序产品 | |
| CN112315463B (zh) | 一种婴幼儿听力测试方法、装置及电子设备 | |
| CN109934142A (zh) | 用于生成视频的特征向量的方法和装置 | |
| CN112242198A (zh) | 基于大数据的失语症个性化治疗方案推荐方法及系统 | |
| CN108460364B (zh) | 用于生成信息的方法和装置 | |
| WO2025201223A1 (zh) | 用于智能辅助交互的电子设备和方法 | |
| CN115641938A (zh) | 一种基于前庭性偏头痛康复训练的虚拟现实方法及系统 | |
| JP6927540B1 (ja) | 情報処理装置、情報処理システム、情報処理方法及びプログラム | |
| CN110348406B (zh) | 参数推断方法及装置 | |
| CN111062995B (zh) | 生成人脸图像的方法、装置、电子设备和计算机可读介质 | |
| CN110545386B (zh) | 用于拍摄图像的方法和设备 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 25775057 Country of ref document: EP Kind code of ref document: A1 |