WO2026016142A1 - 用于执行用户任务的方法、装置、设备和介质 - Google Patents
用于执行用户任务的方法、装置、设备和介质Info
- Publication number
- WO2026016142A1 WO2026016142A1 PCT/CN2024/106257 CN2024106257W WO2026016142A1 WO 2026016142 A1 WO2026016142 A1 WO 2026016142A1 CN 2024106257 W CN2024106257 W CN 2024106257W WO 2026016142 A1 WO2026016142 A1 WO 2026016142A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- user
- candidate
- image
- disclosure
- task
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Pending
Links
Classifications
-
- B—PERFORMING OPERATIONS; TRANSPORTING
- B25—HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
- B25J—MANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
- B25J9/00—Program-controlled manipulators
- B25J9/16—Program controls
Definitions
- the exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks.
- Robotics technology has developed rapidly and is widely used in many technological fields.
- Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging.
- robots In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed.
- robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
- a method for performing a user task is provided.
- a user task is received from a user, instructing a robotic device to acquire a first object.
- At least one candidate method for processing the first object is determined.
- a first message instructing the user of the at least one candidate method is provided.
- the robotic device acquires the first object.
- an apparatus for performing a user task includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module configured to determine at least one candidate method for processing the first object; a message providing module configured to provide the user with a first message indicating the at least one candidate method; and an acquiring module configured to, based on the user's selection of a candidate method from the at least one candidate method, cause the robot to...
- the device acquires the first object.
- an electronic device in a third aspect of this disclosure, includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
- a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
- a computer program product comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
- Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure
- Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks
- Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure
- Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure
- Figure 5 shows a block diagram of a process for identifying objects from an image according to some implementations of this disclosure
- Figure 6 shows a block diagram of the process of invoking the action model according to some implementations of this disclosure
- Figure 7 shows a schematic diagram of an example interface according to some implementations of this disclosure.
- Figure 8 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure
- Figure 9 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure.
- Figure 10 shows a block diagram of a device capable of implementing various implementations of the present disclosure.
- the term “comprising” and similar terms should be understood as open inclusion, i.e., “including but not limited to”.
- the term “based on” should be understood as “at least partially based on”.
- the term “one implementation” or “the implementation” should be understood as “at least one implementation”.
- the term “some implementations” should be understood as “at least some implementations”.
- Other explicit and implicit definitions may also be included below.
- the term “model” can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and/or future-developed technical solutions.
- a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information.
- This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
- a prompt message in response to a user's active request, can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format.
- the pop-up window can also include a selection control allowing the user to choose whether to "agree” or "disagree” to provide personal information to the electronic device.
- the term "in response to” as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
- robots and machine learning technologies have been widely applied in various scenarios.
- robots typically can only perform pre-set, fixed tasks and cannot perform different user tasks according to user needs.
- robotic devices struggle to perform tasks in a manner appropriate to the user's current requirements.
- step 160 the robot performs the desired task according to the user's preferred operating method. At this point, the robot's operation can be effectively controlled to execute the desired task.
- Figure 1 shows a block diagram 100 of the application environment according to an exemplary implementation of this disclosure.
- a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks.
- the physical space 160 can include, but is not limited to, one or more rooms.
- the physical space 160 can include, but is not limited to, a living room, bedroom, study, kitchen, toilet, etc., or a combination of one or more of the above.
- the robot device 110 may include multiple parts.
- the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device.
- the user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks.
- the robot device 110 may include an arm 113 for performing actions such as grasping and releasing.
- the arm 113 can grasp an object and move it to a desired position, and so on.
- the robot device 110 may also include a data acquisition unit 114.
- the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc.
- the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc.
- the robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
- the physical environment 160 may include one or more acquisition units 130, ..., and 132.
- acquisition units 130 For example, one or more image acquisition devices may be deployed in a room to acquire images of the room from various angles.
- the physical environment 160 may include a control device 140, which can control one or more acquisition units 130, ..., via a network (not shown). And 132, etc.
- control device 140 can control various electrical devices in physical space 160.
- a machine learning model (e.g., model 150) may be provided to manage physical space 160.
- model 150 may be located inside physical space 160, alternatively and/or additionally, model 150 may be located at a remote device outside physical space 160, and control device 140, robotic device 110, or other device may access the remote model 150 via a network.
- Model 150 may include one or more models. If model 150 includes multiple models, these multiple models may include multiple types of models. Model 150 may, for example, include at least a language model (LM) and an action model. The language model, by learning from a large corpus, is capable of question answering. The action model can control the robotic device 110 to perform various actions. Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
- LM language model
- the action model can control the robotic device 110 to perform various actions.
- Model 150 may also include, for example, an image recognition model, a text recognition model, and so on.
- user 120 can instruct robot device 110 to manipulate various objects in physical space 110.
- objects can be various items in the home environment.
- user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
- Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
- the robot device 110 in physical space 160 can receive a user task 210 from user 120.
- user task 210 can instruct robot device 110 to obtain a first object (for ease of description, the first object can be referred to as the target object).
- the first object can be referred to as the target object.
- user 120 can say "I want to eat an apple" in natural language; in the example of Figure 2, the first object is "apple” (e.g., first object 220).
- Robot device 110 can... Acquire a spatial image of the physical space 160 where the robot is located.
- the spatial image can be acquired via at least one of acquisition units 114, 130, ..., and 132.
- At least one candidate method for processing the first object can be determined.
- a user might have multiple ways of consuming an apple, such as eating it directly, peeling it, or juicing it, thus the method the user might prefer can be determined.
- the robot device 110 can be instructed to move and acquire the first object 220.
- the methods described above can be executed on any computing device with computing capabilities.
- the methods described above can be executed using an application deployed on robot device 110.
- an application can be deployed on control device 140 to execute the methods described above.
- the powerful processing capabilities of model 150 can be invoked to determine at least one candidate method for processing the first object 220.
- robot device 110 can provide user 120 with a first message indicating at least one candidate method.
- the first message can be displayed to user 120 via interaction unit 112.
- robot device 110 can acquire the first object.
- an apple can be directly provided to the user, an apple and a fruit knife can be provided, or apple juice can be provided, etc.
- a robotic device can perform user tasks in a variety of different operating modes.
- the robotic device can provide the user with at least one candidate mode and, based on the user's selection, process the target object in the manner the user desires. This improves the flexibility and accuracy of the robotic device in performing tasks under different user needs, thereby completing the expected user task.
- the spatial image can be derived from at least one of the following: Image acquisition devices at the robot device and in the physical space.
- Figure 3 shows a block diagram 300 of an image acquisition process according to some implementations of this disclosure.
- spatial images e.g., one or more images 310) of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images of various locations in the physical space, thereby facilitating the search for target objects.
- spatial images of the physical space 160 can be acquired from acquisition units 130, ..., and 132.
- acquisition units 130, ..., and 132 can be pre-deployed at designated locations within the physical space 160, such as a corner of the ceiling, etc.
- images of the physical space 160 taken from a top-down perspective can be obtained, facilitating an overall understanding of the layout of the physical space 160 and thus aiding in the location of target objects.
- whether a spatial image includes a first object can be determined in various ways.
- the first object can be identified from the spatial image based on image recognition technology.
- a prompt word can be constructed and input into the model to invoke the model's processing capabilities to identify the first object from the spatial image.
- the prompt word could be, for example, "Please identify 'apple' from the following image," and the acquired image and prompt word can be submitted to the model.
- the model can process images, and if the image includes a first object, the model can output the location of the object (e.g., the coordinates of the object's region in the image, and/or directly output the image of the region where the object is located, etc.).
- the location of a target object in an image can be detected in various ways so that a robotic device can subsequently acquire that target object.
- At least one candidate method for processing the first object is determined. For example, in a real-world scenario, a user may have multiple ways of consuming an apple, such as eating it directly, peeling it, or juicing it.
- a first prompt word can be obtained based on the user task and the user's user information; and at least one candidate method can be determined based on the first response of a machine learning model to the first prompt word.
- Figure 4 shows a block diagram 400 of the process of invoking a language model according to some implementations of this disclosure.
- the first object 220 to be obtained can be determined from the user task 210 as "apple”.
- a first prompt word 420 can be obtained based on the user task 210 and the user information 410.
- the first prompt word 420 can be expressed, for example, as: "Please determine the method of eating 'apple' according to the user information", or as: “According to the user information, how might 'apple' be eaten", and so on.
- user information 410 may include the user's past methods of handling the first object or methods pre-set by the user for handling the first object, such as the user usually preferring to "eat it directly".
- the first cue word 420 can be input into the language model 430 to determine multiple ways to eat the apple. Alternatively and/or additionally, if the user is found to frequently "peel and eat” based on usage history, then "peel and eat” can be determined as a candidate method.
- the language model 430 is a trained and fine-tuned model with rich knowledge of performing tasks across multiple domains.
- the language model 430 can determine the first object 220 as "apple” from the user task 210, and determine the methods of eating the "apple".
- Figure 4 is merely illustrative; the language model 430 can process one or more user information simultaneously and determine one or more candidate methods for processing the first object. For example, assuming the user likes both peeling and juicing apples, "peeling" and “juicing” can be identified as candidate methods.
- the model can determine "banana” as the first object and can determine multiple methods of eating bananas.
- the powerful processing capabilities and rich knowledge of the model can be leveraged to determine at least one candidate method for processing the target object. Subsequently, the robotic device can be instructed to acquire the target object based on the candidate method selected by the user.
- the technical solution of this disclosure can process the target object in the way currently desired by the user. In this way, the convenience and accuracy of task execution can be improved, and the expected user actions can be completed in a richer and more diverse manner. Task.
- Figure 5 illustrates a block diagram 500 of a process for identifying objects from an image according to some implementations of this disclosure.
- objects 510 (apple) and 520 (fruit knife) can be identified from image 310.
- an image of the physical space where the robotic device is located can be acquired, and a candidate method can be determined in response to an image indicating that the physical space includes a second object associated with a candidate method in at least one candidate method.
- first prompt words can be obtained based on images. Specifically, in the process of determining the first object, multiple second objects can be identified from the spatial image, and these second objects are associated with the processing of the first object. At least one candidate method is then determined based on each of the multiple second objects. Leveraging the powerful processing capabilities and rich knowledge of the model, first prompt words for determining at least one candidate method can be obtained based on the image.
- the knowledge base can be predefined and includes the relationships between the first and second objects.
- the knowledge base could include: a fruit knife can be used to peel an apple, a juicer can be used to juice an apple, and so on.
- image recognition and/or machine learning techniques can be used to determine various objects in physical space, thereby identifying at least one candidate method that may be used to process the first object. In this way, the diversity of methods for processing the first object can be improved.
- a second object corresponding to a candidate object can be determined based on an image of the physical space where the robot device is located.
- the robot device is instructed to acquire both the first and second objects.
- object 510 "apple” can be identified as the first object
- object 520 "fruit knife” as the second object; both are movable objects.
- the robot device can use its arm to grasp the apple and the fruit knife for peeling.
- the robot device can provide the user with the apple and the fruit knife so that the user can determine whether to peel it based on their own needs.
- the robot in response to determining that the second object is a non-movable object, can be instructed to...
- the system can utilize a second object to process a first object and instruct a robotic device to acquire the processed first object.
- the robotic device can move to the location of the second object and directly process the first object using the second object (e.g., putting an apple into the juicer to extract juice) to obtain the processed first object (e.g., apple juice).
- different processing methods can be determined by identifying a movable or immovable second object, improving the flexibility of task execution.
- the operation method for manipulating the second object can be determined based on an image and user tasks, and the robotic device can be instructed to operate the second object according to the operation method in order to process the first object.
- the handle of a fruit knife can be identified from an image, at which point it can be determined that the handle needs to be grasped and the blade moved to peel the apple.
- a second cue word can be constructed based on the image and the user task, and the model can be queried about how to perform the user task using a second object in the image.
- the second cue word could be, for example, "Determine how to use a fruit knife from the following images," and the second cue word and the corresponding image can be sent to the model.
- the model can then return: Hold the handle and move the blade. Subsequently, the robot can be instructed to hold the handle and move the blade to peel the apple.
- a series of instructions for performing the peeling action can be transmitted to the robot to control it.
- the powerful processing capabilities of the model can be invoked to solve unknown problems in complex environments, thereby determining the actions that the robot needs to perform. In this way, the robot's ability to handle complex tasks can be improved, thus performing the user task more accurately.
- a motion model can be used to determine the specific actions to be performed by the robot device.
- Figure 6 shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure.
- a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device.
- This motion model can be a pre-trained and fine-tuned model.
- the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm.
- POS1, (7) the positions of the robot arm's joints
- the state of the tool e.g., a gripper, a cutting tool, etc.
- Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state.
- the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
- An instruction 610 (e.g., "grab an apple") can be input to the motion model 630.
- the instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model.
- the current state of the robot device can be obtained, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, speed, acceleration, etc., of each joint in the arm, and/or the wheels and/or other movable devices of the robot device at the next time point can be determined.
- the determined action 640 can be used to control the state of the robot device at the next time point.
- instruction 610 can be used to perform other complex operations, such as “peeling an apple,” “operating a juicer,” and so on.
- the current image and the current posture of the robot device can be continuously acquired to drive the robot device to the next posture.
- a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.
- the robot can be controlled in environments such as Chinese, English, Japanese, and French. Specifically, it can be based on machine learning techniques.
- the robot provides multilingual capabilities to control it in application environments using different languages.
- eating an apple above describes the process of using a robotic device to perform a user task
- the robotic device can be controlled to perform other user tasks, such as retrieving other items from a room, processing an item in multiple other ways, etc.
- a user instructs the robotic device to retrieve a slice of bread; at least one candidate method could include: eating it directly, toasting it with a toaster, etc.
- At least one candidate method is determined based on the first response to the first prompt word using a machine learning model. See Figure 7 for further details; Figure 7 shows a schematic diagram 700 of an example interface according to some implementations of this disclosure. As shown in Figure 7, the selection of candidate methods can be achieved through an interaction unit 112 on the robot device 110.
- the interaction unit 112 can be a display screen with interactive functionality, and the content displayed on the screen can include multiple controls, such as a first control 710, a second control 720, and a third control 730.
- Each of the multiple controls can correspond to a candidate method; for example, the first control 710 corresponds to "eat directly,” the second control 720 corresponds to “peel and eat,” and the third control 730 corresponds to "juice and eat.”
- the user can select the processing method for the first object by triggering the controls. Utilizing some implementations of this disclosure can greatly enrich the interaction content between the user and the robot device, expanding the diversity of task execution.
- a user can interact with the robot device via the interaction unit 112, for example, the user inputs a task represented by text and/or images, and controls the robot device to perform the task.
- the user can specify the execution conditions of the task, for example, to execute the task immediately, to execute the task after a predetermined time, or to execute the task when a predetermined condition is determined to be met (e.g., when the user makes the action of eating an apple), etc.
- the robotic device can provide the user with various messages. For example, assuming the robotic device cannot find other objects that can be used to process the apple, it can ask the user if they want to eat it directly. Or, assuming the robotic device has not found any apples but only tools for processing them, it can inform the user that no apples were found, and so on. Alternatively and/or additionally, the robotic device can ask the user where the desired object can be found and then proceed to the user-specified location to search for the desired object. Alternative and/or additional locations, if the desired object cannot be found, the robotic device can ask the user whether they need to purchase, etc.
- various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment.
- a Global Positioning System GPS
- satellite signals can be used to determine the precise position of the robotic device.
- a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network.
- a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point.
- the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices.
- An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
- Alternate and/or additional locations can be determined using a visual positioning system to pinpoint the location of the robotic device and/or individual objects.
- a map of the physical space can be pre-acquired, and the locations of each object can be marked on this map.
- the robotic device can utilize echo detection units to detect distances to surrounding objects and, by combining the acquired images with the physical space map, determine the precise location of each object.
- CAD computer-aided design
- GIS geographic information systems
- tracking units can be added to remote controls for household appliances (e.g., television remotes, air conditioner remotes) so that the robotic device can promptly acquire the precise location of important objects, and so on.
- household appliances e.g., television remotes, air conditioner remotes
- the robot's initial position and desired destination can be determined based on the methods described above.
- the robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
- the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location.
- Constraints i.e., the constraints that should be followed during task execution, can be determined using a language model and/or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model.
- prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX,” or "Please determine the precautions during the movement of object XXX,” etc.
- an object e.g., bottled water, plate, bowl, etc.
- the object's original posture should be maintained (e.g., remaining vertical and not tilted).
- constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met.
- safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
- users can select different candidate methods for performing tasks, and the robotic device can complete the expected user tasks in a wider variety of ways. In this way, the flexibility and accuracy of the robotic device in performing tasks under different needs can be improved, thereby completing the expected user tasks.
- Figure 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of this disclosure.
- a user task is received from a user, instructing a robotic device to acquire a first object.
- at least one candidate method for processing the first object is determined.
- a first message instructing the user on the at least one candidate method is provided.
- the robotic device acquires the first object.
- determining at least one candidate method includes: obtaining a first prompt word based on the user task and the user's user information, wherein the first prompt word is used to determine... At least one candidate method; and based on the first response to the first prompt word using a machine learning model, at least one candidate method is determined.
- the method 800 further includes: acquiring an image of the physical space where the robot device is located; and determining a candidate mode in response to the image indicating that the physical space includes a second object associated with a candidate mode among at least one candidate mode.
- obtaining the first prompt word further includes: obtaining an image of the physical space where the robot device is located; and obtaining the first prompt word based on the image.
- obtaining the first object includes: determining a second object corresponding to a candidate method based on an image of the physical space where the robot device is located; and in response to determining that the second object is a movable object, the robot device obtains the first object and the second object.
- obtaining the first object includes: in response to determining that the second object is an immovable object, the robotic device uses the second object to process the first object; and the robotic device obtains the first object to be processed.
- using a second object to process a first object includes: determining an operation mode for operating the second object based on an image and a user task; and a robotic device operating the second object according to the operation mode in order to process the first object.
- determining the operation mode for operating the second object includes: obtaining a second prompt word based on an image and a user task, the second prompt word being used to determine the operation mode by which the second object in the image performs the user task; and determining at least one candidate mode based on a first response of a machine learning model to the first prompt word.
- FIG. 9 shows a block diagram of an apparatus 900 for performing a user task according to some implementations of the present disclosure.
- the apparatus 900 includes: a receiving module 910 configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module 920 configured to determine at least one candidate method for processing the first object; and a message providing module 930 configured to provide the user with instructions indicating the at least one candidate method. First message; and acquisition module 940, configured to enable the robotic device to acquire the first object based on the user's selection of a candidate method among at least one candidate method.
- the determining module 920 is further configured to: obtain a first prompt word based on the user task and the user's user information, the first prompt word being used to determine at least one candidate method; and determine at least one candidate method based on a first response of a machine learning model to the first prompt word.
- the acquisition module 940 is further configured to: acquire an image of the physical space where the robot device is located; the determination module 920 is further configured to: determine a candidate mode in response to an image indicating that the physical space includes a second object associated with a candidate mode among at least one candidate mode.
- the acquisition module 940 is further configured to: acquire an image of the physical space where the robot device is located; the determination module 920 is further configured to: acquire a first prompt word based on the image.
- the acquisition module 940 is further configured to: determine a second object corresponding to a candidate method based on an image of the physical space where the robot device is located; and in response to determining that the second object is a movable object, cause the robot device to acquire the first object and the second object.
- the acquisition module 940 is further configured to: in response to determining that the second object is an immovable object, cause the robot device to process the first object using the second object; and cause the robot device to acquire the first object to be processed.
- the acquisition module 940 is further configured to: determine an operation mode for operating the second object based on the image and the user task; and cause the robot device to operate the second object in accordance with the operation mode in order to process the first object.
- the determining module 920 is further configured to: obtain a second prompt word based on an image and a user task, the second prompt word being used to determine the operation mode by which a second object in the image performs the user task; and determine at least one candidate mode based on a first response of a machine learning model to the first prompt word.
- Figure 10 shows a block diagram of a device 1000 capable of implementing various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementation described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above.
- the computing device 1000 is in the form of a general-purpose computing device.
- Components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060.
- the processing unit 1010 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 1000.
- Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media.
- Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
- Storage device 1030 can be removable or non-removable media and may include machine-readable media, such as flash drives, disks, or any other media capable of storing information and/or data (e.g., training data for training) and accessible within computing device 1000.
- the computing device 1000 may further include additional removable/non-removable, volatile/non-volatile storage media.
- disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
- each drive may be connected to a bus (not shown) via one or more data media interfaces.
- the memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.
- the communication unit 1040 enables communication with other computing devices via a communication medium. (See attached image) Furthermore, the functionality of the components of computing device 1000 can be implemented as a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
- PCs network personal computers
- Input device 1050 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
- Output device 1060 can be one or more output devices, such as a monitor, speaker, printer, etc.
- Computing device 1000 can also communicate with one or more external devices (not shown) via communication unit 1040 as needed. These external devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with computing device 1000, or with any device (e.g., network card, modem, etc.) that enables computing device 1000 to communicate with one or more other computing devices. Such communication can be performed via input/output (I/O) interface (not shown).
- I/O input/output
- a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
- a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
- a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
- These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, these instructions produce the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
- These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, programmable data processing device, and/or other device to operate in a particular manner.
- the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of a flowchart and/or block diagram.
- Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions/actions specified in one or more boxes of a flowchart and/or block diagram.
- each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
- the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
- each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.
Landscapes
- Engineering & Computer Science (AREA)
- Robotics (AREA)
- Mechanical Engineering (AREA)
- Manipulator (AREA)
Abstract
一种用于执行用户任务的方法、装置、设备和介质。其中,用于执行用户任务的方法包括如下步骤:接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定用于处理第一对象的至少一个候选方式;向用户提供指示至少一个候选方式的第一消息;基于用户针对至少一个候选方式中的候选方式的选择,机器人设备获取第一对象。通过该方法,用户可以选择不同的用于执行任务的候选方式,以更多样的形式完成预期的用户任务。
Description
本公开的示例性实现方式总体涉及机器人领域,特别地涉及利用机器人来执行用户任务的方法、装置、设备和计算机可读存储介质。
机器人技术已经得到了迅速发展,并且已经被广泛地用于多个技术领域。目前已经开发出了多种专用机器人设备,例如,在工业环境中,可以使用机器人来执行加工、抓取、分类、包装等多种任务。又例如,在家居环境下,已经开发出了扫地机器人、擦玻璃机器人,等等。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。
发明内容
在本公开的第一方面,提供了一种用于执行用户任务的方法。在该方法中,接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象。确定用于处理第一对象的至少一个候选方式。向用户提供指示至少一个候选方式的第一消息。基于用户针对至少一个候选方式中的候选方式的选择,机器人设备获取第一对象。
在本公开的第二方面,提供了一种用于执行用户任务的装置。该装置包括:接收模块,被配置用于接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定模块,被配置用于确定用于处理第一对象的至少一个候选方式;消息提供模块,被配置用于向用户提供指示至少一个候选方式的第一消息;以及获取模块,被配置用于基于用户针对至少一个候选方式中的候选方式的选择,使得机器人
设备获取第一对象。
在本公开的第三方面,提供了一种电子设备。该电子设备包括:至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令,指令在由至少一个处理单元执行时使电子设备执行根据本公开第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,计算机程序在被处理器执行时使处理器实现根据本公开第一方面的方法。
在本公开的第五方面,提供了一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据本公开第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实现方式的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
在下文中,结合附图并参考以下详细说明,本公开各实现方式的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标注表示相同或相似的元素,其中:
图1示出了根据本公开的一个示例性实现方式的应用环境的框图;
图2示出了根据本公开的一些实现方式的用于执行用户任务的框图;
图3示出了根据本公开的一些实现方式的图像采集过程的框图;
图4示出了根据本公开的一些实现方式的调用语言模型的过程的框图;
图5示出了根据本公开的一些实现方式的从图像中识别对象的过程的框图;
图6示出了根据本公开的一些实现方式的调用动作模型的过程的框图;
图7示出了根据本公开的一些实现方式的示例界面的示意图;
图8示出了根据本公开的一些实现方式的用于执行用户任务的方法的流程图;
图9示出了根据本公开的一些实现方式的用于执行用户任务的装置的框图;以及
图10示出了能够实施本公开的多个实现方式的设备的框图。
下面将参照附图更详细地描述本公开的实现方式。虽然附图中示出了本公开的某些实现方式,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实现方式,相反,提供这些实现方式是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实现方式仅用于示例性作用,并非用于限制本公开的保护范围。
在本公开的实现方式的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实现方式”或“该实现方式”应当理解为“至少一个实现方式”。术语“一些实现方式”应当理解为“至少一些实现方式”。下文还可能包括其他明确的和隐含的定义。如本文中所使用的,术语“模型”可以表示各个数据之间的关联关系。例如,可以基于目前已知的和/或将在未来开发的多种技术方案来获取上述关联关系。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当根据相关法律法规通过适当的方式对本公开所涉及个人信息的
类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限制性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式,例如可以是弹出窗口的方式,弹出窗口中可以以文字的方式呈现提示信息。此外,弹出窗口中还可以承载供用户选择“同意”或“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其他满足相关法律法规的方式也可应用于本公开的实现方式中。
在此使用的术语“响应于”表示相应的事件发生或者条件得以满足的状态。将会理解,响应于该事件或者条件而被执行的后续动作的执行时机,与该事件发生或者条件成立的时间,二者之间未必是强关联的。例如,在某些情况下,后续动作可在事件发生或者条件成立时立即被执行;而在另一些情况下,后续动作可在事件发生或者条件成立后经过一段时间才被执行。
示例环境
近年来,机器人技术和机器学习技术已经被广泛应用于多个应用场景。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。尤其是,机器人设备难以根据用户当前需要的方式执行相应的任务。
目前已经开发了执行特定任务的简单机器人设备,然而,此类简单机器人设备并不能理解复杂的用户指令,也不能在复杂的物理空间
160中按照用户希望的操作方式来执行期望的任务。此时,期望可以以有效的方式控制机器人的操作,进而执行期望的任务。
根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图1描述根据本公开的一个示例实现方式的应用环境,图1示出了根据本公开的一个示例性实现方式的应用环境的框图100。如图1所示,机器人设备110和用户120可以位于物理空间160中,并且用户120可以控制机器人设备110来执行多种任务。物理空间160可以包括但不限于一个或者多个房间。例如,在家居环境下,物理空间160可以包括但不限于客厅、卧室、书房、厨房、厕所,等等,或者包括以上一个或者多个的组合。
如图1所示,机器人设备110可以包括多个部分。例如,控制单元111可以作为机器人设备110的控制中心,可以向控制单元111中加载应用程序,以便控制机器人设备中的各个部分。用户120可以使用交互单元112来与机器人设备110交互,例如,向机器人设备110输入控制指令,以便利用机器人设备110执行期望的任务。机器人设备110可以包括手臂113,用于执行抓取、释放等动作。例如,手臂113可以抓取某个对象,并且将该对象移动至期望的位置,等等。
备选地和/或附加地,机器人设备110还可以包括采集单元114。在此,采集单元114可以包括多种类型,例如,图像采集单元、声音采集单元,等等。备选地和/或附加地,机器人设备110可以进一步包括用于检测周围物体的感测单元,例如,可以基于激光来检测机器人与周围物体的距离,等等。机器人设备110还可以包括驱动单元115,例如,机器人设备110可以被部署在可移动的基座之上,并且驱动单元115可以驱动基座的轮子来按照期望的路径移动。
物理环境160可以包括一个或者多个采集单元130、…、以及132,例如,可以在房间中部署一个或者多个图像采集设备以便从各个角度采集房间的图像。物理环境160可以包括控制设备140,该控制设备140可以经由网络(未示出)控制一个或者多个采集单元130、…、
以及132,等等。备选地和/或附加地,在智能家居环境下,控制设备140可以控制物理空间160中的各种电器设备。
备选地和/或附加地,可以提供机器学习模型(例如,模型150)来管理物理空间160。应当理解,尽管图1示出了模型150位于物理空间160内部,备选地和/或附加地,该模型150可以位于物理空间160之外的远程设备处,并且控制设备140、机器人设备110或者其他设备可以经由网络来访问远程的模型150。
模型150可以包括一个或多个模型。如果模型150包括多个模型,这多个模型可以包括多个类型的模型。模型150例如可以至少包括语言模型(LM)和动作模型。语言模型通过从大量语料中学习,能够具备问答能力。动作模型可以控制机器人设备110来执行各种动作。模型150例如还可以包括图像识别模型、文本识别模型等等。
如图1所示,用户120可以指示机器人设备110来操作物理空间110中的各种对象。在此,对象可以是家居环境中的各种物品,例如,用户120可以指令机器人设备110来在物理空间160中寻找某个对象;又例如,用户120可以指令机器人设备110来将找到对象放置到指定位置,等等。
执行任务的概要
为了至少部分地解决现有技术中的不足,根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图2描述根据本公开的一个示例性实现方式的概要,该图2示出了根据本公开的一些实现方式的用于执行用户任务的框图200。
如图2所示,物理空间160中的机器人设备110可以接收来自用户120的用户任务210。此时,用户任务210可以指示机器人设备110来获取第一对象(为了便于描述可以将第一对象称为目标对象)。例如,用户120可以以自然语言说出“我要吃苹果”,在图2的示例中第一对象为“苹果”(例如,第一对象220)。机器人设备110可以
获取机器人设备所在的物理空间160的空间图像。例如,可以经由采集单元114、130、…、以及132中的至少任一项来获取空间图像。
继而,可以确定用于处理第一对象的至少一个候选方式。例如在实际场景中,用户针对苹果可能会有直接吃、削皮吃、榨汁吃等多种食用方法,因而可以确定用户可能想要的食用方法。如图2所示,假设房间中包括第一对象220,则可以指示机器人设备110来移动并获取第一对象220。
根据本公开的一些实现方式,可以在具有计算能力的任何计算设备处执行上文描述的方法。例如,可以利用部署在机器人设备110处的应用程序来执行上文描述的方法。备选地和/或附加地,可以在控制设备140处部署应用程序,以便执行上述方法。具体地,可以调用模型150的强大处理能力,以便确定用于处理第一对象220的至少一个候选方式。继而,机器人设备110可以向用户120提供指示至少一个候选方式的第一消息。第一消息可以经由交互单元112向用户120展示。基于用户120针对至少一个候选方式中的候选方式的选择,机器人设备110可以获取第一对象。例如,可以直接向用户提供苹果、可以提供苹果和水果刀、或者可以提供榨汁后的苹果汁,等等。
利用本公开的示例性实现方式,机器人设备可以以多种不同的操作方式执行用户任务。以此方式,机器人设备可以向用户提供至少一个候选方式,并基于用户的选择,以用户期望的方式处理目标对象。以此方式,可以提高机器人设备在不同的用户需求下执行任务的灵活度和精确度,进而完成预期的用户任务。
执行任务的详细过程
已经描述了根据本公开的一些实现方式的概要,在下文中,将描述有关执行用户任务的更多细节。为了便于描述,在下文中仅以控制机器人设备110处理苹果作为示例,来描述执行用户任务的更多细节。
根据本公开的一些实现方式,空间图像可以来自以下至少任一项:
机器人设备处的采集设备和物理空间中的采集设备。参见图3描述图像采集的更多细节,该图3示出了根据本公开的一些实现方式的图像采集过程的框图300。如图3所示,可以从机器人设备110处的采集单元114获取物理空间160的空间图像(例如,一个或者多个图像310)。由于机器人设备110可以在物理空间160中自由移动,因而采集单元114可以采集物理空间中的各个位置的图像,进而便于寻找目标对象。
备选地和/或附加地,可以从采集单元130、…、以及132来获取物理空间160的空间图像。在此,采集单元130、…、以及132可以被预先部署在物理空间160内的指定位置,例如,天花板的拐角位置,等等。以此方式,可以获得以俯视角度拍摄的物理空间160的图像,进而便于从整体上了解物理空间160的布局,进而便于定位目标对象。
根据本公开的一些实现方式,可以基于多种方式来确定空间图像是否包括第一对象。例如,可以基于图像识别技术来从空间图像中识别第一对象。备选地和/或附加地,可以构建提示词并且向模型输入该提示词,以便调用模型的处理能力来从空间图像中识别第一对象。提示词例如可以表示为:“请从如下图像中识别‘苹果’”,并且向模型提交采集的图像和提示词。
模型可以处理图像,并且在图像包括第一对象的情况下,模型可以输出对象所在的位置(例如,对象在图像中的区域坐标,和/或直接输出对象所在区域的图像,等等)。利用本公开的一些实现方式,可以基于多种方式检测图像中目标对象的位置,以便机器人设备后续获取该目标对象。
根据本公开的一些实现方式,确定用于处理第一对象的至少一个候选方式。例如在实际场景中,用户针对苹果可能会有直接吃、削皮吃、榨汁吃等多种食用方法。具体地,在确定至少一个候选方式的过程中,可以基于用户任务和用户的用户信息,获取第一提示词;以及基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。
参见图4描述有关确定至少一个候选方式的更多细节,图4示出了根据本公开的一些实现方式的调用语言模型的过程的框图400。如图4所示,可以从用户任务210确定将要获取的第一对象220为“苹果”,此时可以基于用户任务210和用户信息410,获取第一提示词420。第一提示词420例如可以表示为:“请根据用户信息确定‘苹果’的食用方法”,又例如,第一提示词420可以表示为:“根据用户信息,‘苹果’可能被怎么食用”,等等。
根据本公开的一些实现方式,用户信息410可以包括用户以往的用于处理第一对象的方式或用户预先设置的用于处理第一对象的方式,例如用户往常更倾向于“直接吃”。第一提示词420能够被输入至语言模型430,以确定苹果的多种食用方法。备选地和/或附加地,如果基于使用历史发现用户经常“削皮吃”,则可以确定候选方式为“削皮吃”。
根据本公开的一些实现方式,语言模型430是经过训练和微调的模型,并且具有执行多个领域的任务的丰富知识。语言模型430可以从用户任务210中确定第一对象220为“苹果”,并且确定“苹果”的食用方法。图4仅仅是示意性的,语言模型430可以同时处理一个或者多个用户信息,并且确定用于处理第一对象的一个或者至少一个候选方式。例如,假设用户既喜欢削皮吃苹果,又喜欢榨汁吃苹果,则可以将“削皮吃”和“榨汁吃”确定为候选方式。备选地和/或附加地,假设用户任务为“我要吃香蕉”,则模型可以确定“香蕉”为第一对象,并且可以确定香蕉的多种食用方法。
利用本公开的一些实现方式,可以利用模型的强大处理能力和丰富知识,确定用于处理目标对象的至少一个候选方式。继而,可以指示机器人设备基于用户选择的候选方式来获取目标对象。相对于仅能通过预设方式来处理目标对象的已有技术方案而言,本公开的技术方案可以通过用户当前期望的方式处理目标对象。以此方式,可以提高任务执行的便捷性和准确性,以更加丰富多样的形式完成预期的用户
任务。
参见图5描述更多细节,图5示出了根据本公开的一些实现方式的从图像中识别对象的过程的框图500。如图5所示,从图像310中可以识别出对象510(苹果)和对象520(水果刀)。根据本公开的一些实现方式,可以获取机器人设备所在的物理空间的图像,并且响应于图像指示物理空间包括与至少一个候选方式中的候选方式相关联的第二对象,确定候选方式。
此外,还可以基于图像获取第一提示词。具体地,在确定第一对象的过程中,可以从空间图像中识别多个第二对象,多个第二对象与第一对象的处理相关联,以及基于多个第二对象分别确定至少一个候选方式。利用模型的强大处理能力和丰富知识可以基于图像获取用于确定至少一个候选方式的第一提示词。在此,知识库可以是预先定义的,并且包括第一对象与第二对象之间的关联关系。例如,知识库可以包括:水果刀可以用于对苹果进行削皮操作、榨汁机可以用于对苹果进行榨汁操作,等等。
利用本公开的一些实现方式,可以利用图像识别技术和/或机器学习技术确定物理空间中的各个对象,进而确定可能用于处理第一对象的至少一个候选方式。以此方式,可以提高处理第一对象的方法的多样性。
根据本公开的一些实现方式,可以基于机器人设备所在的物理空间的图像,确定与候选方式相对应的第二对象,并且响应于确定第二对象为可移动对象,指示机器人设备获取第一对象和第二对象。如图5所示,在图像310中,可以确定对象510“苹果”为第一对象,对象520“水果刀”为第二对象,两者均为可移动对象。此时,机器人设备可以利用手臂抓取苹果和水果刀,以进行削皮操作。备选地和/或附加地,机器人设备可以向用户提供苹果和水果刀,以便用户基于自身需求确定是否削皮。
相反,响应于确定第二对象为不可移动对象,可以指示机器人设
备利用第二对象处理第一对象,并且指示机器人设备获取加工的第一对象。例如,在确定图像310中对象530“榨汁机”为第二对象的情况下,机器人设备可以移动至第二对象的位置,并直接利用第二对象对第一对象进行加工处理(例如将苹果放入榨汁机中进行榨汁),以得到加工的第一对象(例如苹果汁)。利用本公开的一些实现方式,可以通过识别可移动的或不可移动的第二对象来确定不同的处理方式,提高了任务执行的灵活度。
根据本公开的一些实现方式,在利用第二对象来加工第一对象的过程中,可以基于图像和用户任务,确定用于操作第二对象的操作方式,并且指示机器人设备来按照操作方式,操作第二对象以便处理第一对象。具体地,可以从图像中识别出水果刀的刀柄,此时可以确定需要握住刀柄并移动刀刃从而削去苹果皮。
备选地和/或附加地,可以基于图像和用户任务,构建第二提示词并且询问模型由图像中的第二对象来执行用户任务的操作方式。第二提示词例如可以表示为:“从以下图像中确定使用水果刀的方式”,可以向模型发送第二提示词和相应的图像。此时,模型可以返回:握住刀柄并移动刀刃。继而,可以指示机器人设备握住刀柄并移动刀刃从而削去苹果皮。备选地和/或附加地,可以向机器人设备传输用于执行削皮动作的一系列指令,以便控制机器人设备。利用本公开的一些实现方式,可以调用模型的强大处理能力来解决复杂环境下的未知问题,进而确定需要机器人设备执行的动作。以此方式,可以提高机器人设备处理复杂任务的能力,从而以更为准确的方式来执行用户任务。
根据本公开的一些实现方式,可以利用动作模型来确定机器人设备执行的具体动作。参见图6描述更多细节,图6示出了根据本公开的一些实现方式的调用动作模型的过程的框图600。如图6所示,可以提供动作模型630,该动作模型630可以基于机器人设备的当前状态和指令来确定将要由机器人设备执行的具体动作,该动作模型可以是经过预训练和微调的模型。
应当理解,在此当前状态可以包括多个方面的数据,例如,机器人设备的图像、机器人设备的环境图像、机器人手臂的姿态数据(如,机器人手臂的各个关节的位置(POS1,…))、以及被固定在机器人手臂的末端的工具(例如,夹具,刀具,等等)的状态。例如,可以使用0来表示夹具的关闭状态,并且使用1来表示夹具的开放状态。可以将指令和当前状态输入至动作模型630,进而利用动作模型,基于指令和当前状态来确定由机器人设备将要执行的动作。在此,动作可以表示机器人设备的当前姿态与下一姿态之间的差异,以及工具的当前状态与下一状态之间的差异,等等。
可以向动作模型630输入指令610(例如,“抓取苹果”),在此,指令610可以是以自然语言表示的,并且可以从语言模型的应答中确定该指令610。进一步,可以获取机器人设备的当前状态,动作模型630可以基于输入的数据来确定相应的动作640。例如,可以确定手臂中的各个关节、和/或机器人设备的轮子和/或其他可移动装置在下一时间点的朝向、位置、速度、加速度,等等。进一步,可以利用确定的动作640来控制机器人设备在下一时间点的状态。
尽管上文以“抓取苹果”作为示例描述了确定动作的具体过程,备选地和/或附加地,指令610可以用于执行其他复杂的操作,例如,“削苹果皮”、“操作榨汁机”,等等。此时,可以不断地采集当前的图像和机器人设备当前姿态,以便驱动机器人设备到达下一姿态。
利用本公开的一些实现方式,可以在语言模型和动作模型之间建立关联关系,并且将用户最初输入的以自然语言表达的用户任务转换为由机器人设备可执行的具体动作。以此方式,可以精确地控制机器人设备的动作,进而以更高的效率执行用户任务。
应当理解,尽管上文以中文语言环境为示例描述根据本公开的一个示例实现方式。备选地和/或附加地,可以在多种语言环境中执行根据本公开的一个示例实现方式的技术方案。例如,可以在中文、英文、日文、法文等环境中控制机器人。具体地,可以基于机器学习技术所
提供的多语言能力,来在不同语言的应用环境中控制机器人。进一步,尽管上文以吃苹果作为示例描述了利用机器人设备执行用户任务的过程,备选地和/或附加地,可以控制机器人设备来执行其他用户任务,例如,在房间中获取其他物品,将某个物品以其他多种方式处理,等等。假设用户指示机器人设备取回面包片,至少一个候选方式可以包括:直接吃、用吐司机烤后吃,等等。
根据本公开的一些实现方式,基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。参见图7描述更多细节,图7示出了根据本公开的一些实现方式的示例界面的示意图700。如图7所示,候选方式的选择可以通过机器人设备110上的交互单元112实现。交互单元112可以是具有交互功能的显示屏,显示屏上展示的内容可以包括多个控件,例如第一控件710、第二控件720以及第三控件730。多个控件中的每个控件可以对应一个候选方式,例如第一控件710对应“直接吃”,第二控件720对应“削皮吃”,第三控件730对应“榨汁吃”。用户可以通过触发控件来选择针对第一对象的处理方式。利用本公开的一些实现方式,可以极大程度上丰富用户与机器人设备的交互内容,扩展了任务执行的多样性。
备选地和/或附加地,用户可以经由交互单元112来与机器人设备交互,例如,用户输入以文字和/或图像表示的任务,并且控制机器人设备执行该任务。备选地和/或附加地,用户可以指定任务的执行条件,例如,立即执行任务,在预定时间之后执行任务,或者在确定满足预定条件(例如,在用户做出吃苹果的动作)时执行任务,等等。
根据本公开的一些实现方式,机器人设备可以向用户提供多种消息,例如,假设机器人设备未找到可以用来加工苹果的其他对象,可以询问用户是否需要直接食用。又例如,假设机器人设备没有找到苹果,并且只找到用于加工苹果的工具,机器人设备可以告知用户未找到苹果,等等。备选地和/或附加地,机器人设备可以询问用户在哪里可以找到期望的对象,并且前往用户指定的位置来寻找期望的对象。
备选地和/或附加地,如果不能找到期望的对象,机器人设备可以询问用户是否需要购买,等等。
根据本公开的一些实现方式,可以利用多种定位算法来确定机器人设备、以及各个对象在物理环境中的位置。例如,可以在机器人设备处部署全球定位系统(GPS),并且使用卫星信号来确定机器人设备的精确位置。备选地和/或附加地,可以在机器人设备处部署通信单元,借助于该通信单元与基站之间的信号,并利用通信网络来确定机器人设备的位置。备选地和/或附加地,在物理空间中可以部署Wi-Fi接入点,机器人设备处的通信单元可以与Wi-Fi热点交互,以便经由Wi-Fi信号强度和已知Wi-Fi接入点的位置来确定位置。备选地和/或附加地,机器人设备处的通信单元可以支持蓝牙功能,此时可以使用蓝牙信号和已知的蓝牙设备位置来确定附近设备的位置。可以在机器人设备处部署惯性导航系统,并且使用加速度计和陀螺仪来测量和计算设备在空间中的移动和方向,进而确定机器人设备的位置。
备选地和/或附加地,可以使用视觉定位系统来确定机器人设备和/或各个对象的位置。可以预先获取物理空间的地图,并且在该地图中标注各个对象的位置。机器人设备可以利用回波检测单元来检测与周围对象之间的距离,并且结合采集的图像和物理空间的地图,来确定各个对象的具体位置。具体地,可以使用计算机辅助设计(CAD)和地理信息系统(GIS),并且利用定位算法来确定位置。备选地和/或附加地,可以在物理空间中的重要对象处部署跟踪单元,例如,可以在家用电器的遥控器(例如,电视遥控器、空调遥控器)处添加跟踪单元,以便机器人设备可以及时获取重要对象的精确位置,等等。
根据本公开的一些实现方式,可以基于上文描述的方法来确定机器人设备自身的原始位置和期望去往的目的地位置。机器人设备可以确定从原始位置到目的地位置的路径。例如,可以不断获取周围的环境图像,并且在确保躲避障碍的情况下,不断更新该路径,并且使得机器人设备沿着该路径移动至目的地位置。
根据本公开的一些实现方式,在到达目的地位置之后,机器人设备可以执行指定的任务。例如,可以获取指定的对象并且将该对象移动至相应位置。可以利用语言模型和/或知识库来确定约束条件,也即在执行任务期间应当遵循的约束条件。例如,可以获取图像和相应提示词,并且向语言模型输入该图像和提示词,进而从语言模型接收约束条件。例如,可以确定提示词:“请基于如下图像,确定在移动XXX对象期间应当遵循的约束条件”,或者“请确定移动XXX对象期间的注意事项”,等等。
此时,可以确定在移动对象(例如,瓶装水、盘子、碗等)期间,应当保持对象的原始姿态(例如,保持竖直方向,不会被倾斜)。进一步,可以向动作模型输入约束条件,此时,动作模型输出的一系列动作将会在确保约束条件的情况下,执行相应的任务。利用本公开的一些实现方式,可以确保机器人设备操作期间的安全性,从而避免造成意外损坏某个对象,等等。
利用本公开的示例性实现方式,用户可以选择不同的用于执行任务的候选方式,并且机器人设备可以以更加丰富多样的形式完成预期的用户任务。以此方式,可以提高机器人设备在不同需求下执行任务的灵活度和精确度,进而完成预期的用户任务。
示例过程
图8示出了根据本公开的一些实现方式的用于执行用户任务的方法800的流程图。在框810处,接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象。在框820处,确定用于处理第一对象的至少一个候选方式。在框830处,向用户提供指示至少一个候选方式的第一消息。在框840处,基于用户针对至少一个候选方式中的候选方式的选择,机器人设备获取第一对象。
根据本公开的一些实现方式,确定至少一个候选方式包括:基于用户任务和用户的用户信息,获取第一提示词,第一提示词用于确定
至少一个候选方式;以及基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。
根据本公开的一些实现方式,该方法800进一步包括:获取机器人设备所在的物理空间的图像;以及响应于图像指示物理空间包括与至少一个候选方式中的候选方式相关联的第二对象,确定候选方式。
根据本公开的一些实现方式,获取第一提示词进一步包括:获取机器人设备所在的物理空间的图像;以及基于图像获取第一提示词。
根据本公开的一些实现方式,获取第一对象包括:基于机器人设备所在的物理空间的图像,确定与候选方式相对应的第二对象;以及响应于确定第二对象为可移动对象,机器人设备获取第一对象和所述第二对象。
根据本公开的一些实现方式,获取第一对象包括:响应于确定第二对象为不可移动对象,机器人设备利用第二对象处理第一对象;以及机器人设备获取加工的第一对象。
根据本公开的一些实现方式,利用第二对象来加工第一对象包括:基于图像和用户任务,确定用于操作第二对象的操作方式;以及机器人设备来按照操作方式,操作第二对象以便处理第一对象。
根据本公开的一些实现方式,确定用于操作第二对象的操作方式包括:基于图像和用户任务,获取第二提示词,第二提示词用于确定由图像中的第二对象来执行用户任务的操作方式;以及基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。
示例装置和设备
图9示出了根据本公开的一些实现方式的用于执行用户任务的装置900的框图。该装置900包括:接收模块910,被配置用于接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定模块920,被配置用于确定用于处理第一对象的至少一个候选方式;消息提供模块930,被配置用于向用户提供指示至少一个候选方式的
第一消息;以及获取模块940,被配置用于基于用户针对至少一个候选方式中的候选方式的选择,使得机器人设备获取第一对象。
根据本公开的一些实现方式,确定模块920进一步被配置用于:基于用户任务和用户的用户信息,获取第一提示词,第一提示词用于确定至少一个候选方式;以及基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。
根据本公开的一些实现方式,获取模块940进一步被配置用于:获取机器人设备所在的物理空间的图像;确定模块920进一步被配置用于:响应于图像指示物理空间包括与至少一个候选方式中的候选方式相关联的第二对象,确定候选方式。
根据本公开的一些实现方式,获取模块940进一步被配置用于:获取机器人设备所在的物理空间的图像;确定模块920进一步被配置用于:基于图像获取第一提示词。
根据本公开的一些实现方式,获取模块940进一步被配置用于:基于机器人设备所在的物理空间的图像,确定与候选方式相对应的第二对象;以及响应于确定第二对象为可移动对象,使得机器人设备获取第一对象和第二对象。
根据本公开的一些实现方式,获取模块940进一步被配置用于:响应于确定第二对象为不可移动对象,使得机器人设备利用第二对象处理第一对象;以及使得机器人设备获取加工的第一对象。
根据本公开的一些实现方式,获取模块940进一步被配置用于:基于图像和用户任务,确定用于操作第二对象的操作方式;以及使得机器人设备来按照操作方式,操作第二对象以便处理第一对象。
根据本公开的一些实现方式,确定模块920进一步被配置用于:基于图像和用户任务,获取第二提示词,第二提示词用于确定由图像中的第二对象来执行用户任务的操作方式;以及基于机器学习模型针对第一提示词的第一应答,确定至少一个候选方式。
图10示出了能够实施本公开的多个实现方式的设备1000的框图。
应当理解,图10所示出的计算设备1000仅仅是示例性的,而不应当构成对本文所描述的实现方式的功能和范围的任何限制。图10所示出的计算设备1000可以用于实现上文描述的方法。
如图10所示,计算设备1000是通用计算设备的形式。计算设备1000的组件可以包括但不限于一个或多个处理器或处理单元1010、存储器1020、存储设备1030、一个或多个通信单元1040、一个或多个输入设备1050以及一个或多个输出设备1060。处理单元1010可以是实际或虚拟处理器并且能够根据存储器1020中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高计算设备1000的并行处理能力。
计算设备1000通常包括多个计算机存储介质。这样的介质可以是计算设备1000可访问的任何可以获得的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器1020可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备1030可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据(例如用于训练的训练数据)并且可以在计算设备1000内被访问。
计算设备1000可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图10中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器1020可以包括计算机程序产品1025,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实现方式的各种方法或动作。
通信单元1040实现通过通信介质与其他计算设备进行通信。附
加地,计算设备1000的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,计算设备1000可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备1050可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备1060可以是一个或多个输出设备,例如显示器、扬声器、打印机等。计算设备1000还可以根据需要通过通信单元1040与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与计算设备1000交互的设备进行通信,或者与使得计算设备1000与一个或多个其他计算设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,提供了一种计算机程序产品,其上存储有计算机程序,程序被处理器执行时实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作
的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。
Claims (12)
- 一种用于执行用户任务的方法,包括:接收来自用户的用户任务,所述用户任务指示机器人设备来获取第一对象;确定用于处理所述第一对象的至少一个候选方式;向所述用户提供指示所述至少一个候选方式的第一消息;以及基于所述用户针对所述至少一个候选方式中的候选方式的选择,机器人设备获取所述第一对象。
- 根据权利要求1所述的方法,其中确定所述至少一个候选方式包括:基于所述用户任务和所述用户的用户信息,获取第一提示词,所述第一提示词用于确定所述至少一个候选方式;以及基于机器学习模型针对所述第一提示词的第一应答,确定所述至少一个候选方式。
- 根据权利要求2所述的方法,进一步包括:获取所述机器人设备所在的物理空间的图像;以及响应于所述图像指示所述物理空间包括与所述至少一个候选方式中的候选方式相关联的第二对象,确定所述候选方式。
- 根据权利要求2所述的方法,其中获取所述第一提示词进一步包括:获取所述机器人设备所在的物理空间的图像;以及基于所述图像获取所述第一提示词。
- 根据权利要求1所述的方法,其中获取所述第一对象包括:基于所述机器人设备所在的物理空间的图像,确定与所述候选方式相对应的第二对象;以及响应于确定所述第二对象为可移动对象,所述机器人设备获取所述第一对象和所述第二对象。
- 根据权利要求5所述的方法,其中获取所述第一对象包括:响应于确定所述第二对象为不可移动对象,所述机器人设备利用所述第二对象处理所述第一对象;以及所述机器人设备获取加工的所述第一对象。
- 根据权利要求6所述的方法,其中利用所述第二对象来加工所述第一对象包括:基于所述图像和所述用户任务,确定用于操作所述第二对象的操作方式;以及所述机器人设备来按照所述操作方式,操作所述第二对象以便处理所述第一对象。
- 根据权利要求7所述的方法,其中确定用于操作所述第二对象的所述操作方式包括:基于所述图像和所述用户任务,获取第二提示词,所述第二提示词用于确定由所述图像中的所述第二对象来执行所述用户任务的操作方式;以及基于机器学习模型针对所述第一提示词的第一应答,确定所述至少一个候选方式。
- 一种用于执行用户任务的装置,包括:接收模块,被配置用于接收来自用户的用户任务,所述用户任务机器人设备来获取第一对象;确定模块,被配置用于确定用于处理所述第一对象的至少一个候选方式;消息提供模块,被配置用于向所述用户提供指示所述至少一个候选方式的第一消息;以及获取模块,被配置用于基于所述用户针对所述至少一个候选方式中的候选方式的选择,机器人设备获取所述第一对象。
- 一种电子设备,包括:至少一个处理单元;以及至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至8中任一项所述的方法。
- 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序在被处理器执行时使所述处理器实现根据权利要求1至8中任一项所述的方法。
- 一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据权利要求1至8中任一项所述的方法。
Priority Applications (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2024/106257 WO2026016142A1 (zh) | 2024-07-18 | 2024-07-18 | 用于执行用户任务的方法、装置、设备和介质 |
| CN202480003525.2A CN121752396A (zh) | 2024-07-18 | 2024-07-18 | 用于执行用户任务的方法、装置、设备和介质 |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/CN2024/106257 WO2026016142A1 (zh) | 2024-07-18 | 2024-07-18 | 用于执行用户任务的方法、装置、设备和介质 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2026016142A1 true WO2026016142A1 (zh) | 2026-01-22 |
Family
ID=98436631
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/CN2024/106257 Pending WO2026016142A1 (zh) | 2024-07-18 | 2024-07-18 | 用于执行用户任务的方法、装置、设备和介质 |
Country Status (2)
| Country | Link |
|---|---|
| CN (1) | CN121752396A (zh) |
| WO (1) | WO2026016142A1 (zh) |
Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200242964A1 (en) * | 2017-09-18 | 2020-07-30 | Microsoft Technology Licensing, Llc | Providing diet assistance in a session |
| CN111975772A (zh) * | 2020-07-31 | 2020-11-24 | 深圳追一科技有限公司 | 机器人控制方法、装置、电子设备及存储介质 |
| CN114488879A (zh) * | 2021-12-30 | 2022-05-13 | 深圳鹏行智能研究有限公司 | 一种机器人控制方法以及机器人 |
| CN116810770A (zh) * | 2022-08-31 | 2023-09-29 | 南方科技大学 | 机器人任务规划方法、装置及终端设备 |
-
2024
- 2024-07-18 CN CN202480003525.2A patent/CN121752396A/zh active Pending
- 2024-07-18 WO PCT/CN2024/106257 patent/WO2026016142A1/zh active Pending
Patent Citations (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US20200242964A1 (en) * | 2017-09-18 | 2020-07-30 | Microsoft Technology Licensing, Llc | Providing diet assistance in a session |
| CN111975772A (zh) * | 2020-07-31 | 2020-11-24 | 深圳追一科技有限公司 | 机器人控制方法、装置、电子设备及存储介质 |
| CN114488879A (zh) * | 2021-12-30 | 2022-05-13 | 深圳鹏行智能研究有限公司 | 一种机器人控制方法以及机器人 |
| CN116810770A (zh) * | 2022-08-31 | 2023-09-29 | 南方科技大学 | 机器人任务规划方法、装置及终端设备 |
Also Published As
| Publication number | Publication date |
|---|---|
| CN121752396A (zh) | 2026-03-27 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| US10891484B2 (en) | Selectively downloading targeted object recognition modules | |
| US10134293B2 (en) | Systems and methods for autonomous drone navigation | |
| EP3965068B1 (en) | Delegation of object and pose detection | |
| US20250284284A1 (en) | Online authoring of robot autonomy applications | |
| US20250262759A1 (en) | Information processing system, information processing method, and nonvolatile storage medium capable of being read by computer that stores information processing program | |
| EP4547451A1 (en) | Robot control based on natural language instructions and on descriptors of objects that are present in the environment of the robot | |
| US10789543B1 (en) | Functional object-oriented networks for manipulation learning | |
| US12051252B2 (en) | Location discovery | |
| CN118893635B (zh) | 一种机械臂的控制方法以及机械臂系统 | |
| WO2022259600A1 (ja) | 情報処理装置、情報処理システム、および情報処理方法、並びにプログラム | |
| Yenamandra et al. | Towards open-world mobile manipulation in homes: Lessons from the neurips 2023 homerobot open vocabulary mobile manipulation challenge | |
| WO2026016142A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026016139A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026016140A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026016144A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026016133A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026016137A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| Deguchi et al. | Enhanced robot navigation with human geometric instruction | |
| WO2026016134A1 (zh) | 用于执行用户任务的方法、装置、设备和介质 | |
| WO2026055826A1 (zh) | 获取物品方法、装置、设备和介质 | |
| WO2026055824A1 (zh) | 用于确定机器人设备的动作的方法、装置、设备和介质 | |
| Dizet et al. | RoboCup@ Home Education 2020 Best Performance: RoboBreizh, a modular approach | |
| JP2021010995A (ja) | ロボット制御装置及びロボット | |
| US20240153230A1 (en) | Generalized three dimensional multi-object search | |
| CN121713184A (zh) | 用于执行用户任务的方法、装置、设备和介质 |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 24947534 Country of ref document: EP Kind code of ref document: A1 |