WO2026016136A1 - 用于执行用户任务的方法、装置、设备和介质 - Google Patents

用于执行用户任务的方法、装置、设备和介质

Info

Publication number
WO2026016136A1
WO2026016136A1 PCT/CN2024/106250 CN2024106250W WO2026016136A1 WO 2026016136 A1 WO2026016136 A1 WO 2026016136A1 CN 2024106250 W CN2024106250 W CN 2024106250W WO 2026016136 A1 WO2026016136 A1 WO 2026016136A1
Authority
WO
WIPO (PCT)
Prior art keywords
user
response
book
robot device
disclosure
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/106250
Other languages
English (en)
French (fr)
Inventor
吴弘涛
孔涛
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Youzhuju Network Technology Co Ltd
Original Assignee
Beijing Youzhuju Network Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Youzhuju Network Technology Co Ltd filed Critical Beijing Youzhuju Network Technology Co Ltd
Priority to PCT/CN2024/106250 priority Critical patent/WO2026016136A1/zh
Priority to CN202480003519.7A priority patent/CN121712619A/zh
Publication of WO2026016136A1 publication Critical patent/WO2026016136A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • BPERFORMING OPERATIONS; TRANSPORTING
    • B25HAND TOOLS; PORTABLE POWER-DRIVEN TOOLS; MANIPULATORS
    • B25JMANIPULATORS; CHAMBERS PROVIDED WITH MANIPULATION DEVICES
    • B25J9/00Program-controlled manipulators
    • B25J9/16Program controls

Definitions

  • the exemplary implementations of this disclosure generally relate to the field of robotics, and more particularly to methods, apparatus, devices, and computer-readable storage media for using robots to perform user tasks.
  • Robotics technology has developed rapidly and is widely used in many technological fields.
  • Various specialized robotic devices have been developed; for example, in industrial environments, robots can perform a variety of tasks such as processing, grasping, sorting, and packaging.
  • robots In home environments, for instance, robotic vacuum cleaners and window cleaning robots have been developed.
  • robots typically can only perform pre-set, fixed tasks and cannot perform different user-defined tasks according to user needs.
  • a method for performing a user task is provided.
  • a user task is received from a user, the user task instructing a robotic device to acquire a first object; a second object associated with the first object is determined; and in response to a first image indicating that the first physical space contains both the first and second objects, the robotic device is instructed to acquire the first and second objects.
  • an apparatus for performing a user task includes: a receiving module configured to receive a user task from a user, the user task instructing a robotic device to acquire a first object; a determining module configured to determine a second object associated with the first object; and an acquiring module configured to, in response to a first image indicating that a first physical space in which the robotic device is located includes the first object and the second object, instruct the robotic device to acquire the first object and the second object.
  • an electronic device in a third aspect of this disclosure, includes: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions causing the electronic device to perform the method according to a first aspect of this disclosure when executed by the at least one processing unit.
  • a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, causes the processor to implement the method according to a first aspect of this disclosure.
  • a computer program product comprising a computer program that, when executed by a processor, implements the method according to a first aspect of this disclosure.
  • Figure 1 shows a block diagram of an application environment according to an exemplary implementation of the present disclosure
  • Figure 2 shows a block diagram of some implementations of the present disclosure for performing user tasks
  • Figure 3 shows a block diagram of an image acquisition process according to some implementations of this disclosure
  • Figure 4 shows a flowchart of the process of invoking a language model according to some implementations of this disclosure
  • Figure 5 shows a block diagram of the process of invoking the language model according to some other implementations of this disclosure
  • Figure 6 illustrates the process of invoking the action model according to some implementations of this disclosure.
  • Figure 7 shows a block diagram of the process of obtaining an object according to some implementations of this disclosure.
  • Figure 8 shows a flowchart of a method for performing user tasks according to some implementations of this disclosure
  • Figure 9 shows a block diagram of an apparatus for performing user tasks according to some implementations of the present disclosure.
  • Figure 10 shows a block diagram of a device that can implement various implementations of the present disclosure.
  • the term “comprising” and similar terms should be understood as open inclusion, i.e., “including but not limited to”.
  • the term “based on” should be understood as “at least partially based on”.
  • the term “one implementation” or “the implementation” should be understood as “at least one implementation”.
  • the term “some implementations” should be understood as “at least some implementations”.
  • Other explicit and implicit definitions may also be included below.
  • the term “model” can represent the relationships between various data. For example, the aforementioned relationships can be obtained based on various currently known and/or future-developed technical solutions.
  • a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information.
  • This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
  • a prompt message in response to a user's active request, can be sent to the user, for example, via a pop-up window, where the prompt message can be presented in text format.
  • the pop-up window can also include a selection control allowing the user to choose whether to "agree” or "disagree” to provide personal information to the electronic device.
  • the term "in response to” as used herein refers to a state in which a corresponding event occurs or a condition is satisfied. It will be understood that the timing of subsequent actions performed in response to such event or condition is not necessarily strongly correlated with the time when the event occurs or the condition is met. For example, in some cases, subsequent actions may be performed immediately upon the occurrence of the event or the fulfillment of the condition; while in others, they may be performed some time after the occurrence of the event or the fulfillment of the condition.
  • robots and machine learning technologies have been widely applied in various scenarios.
  • robots typically can only perform pre-set, fixed tasks and cannot execute different user tasks according to user needs.
  • robotic devices struggle to determine user requirements and thus perform corresponding tasks.
  • Simple robotic devices have been developed to perform specific tasks; however, these devices cannot understand complex user instructions, nor can they execute desired tasks according to user commands within complex physical spaces. In this case, it is necessary to efficiently control... To control the operation of the robot and thus perform the desired task.
  • Figure 1 shows a block diagram 100 of an application environment according to an exemplary implementation of this disclosure.
  • a robot device 110 and a user 120 can be located in a physical space 160, and the user 120 can control the robot device 110 to perform various tasks.
  • the physical space 160 can include, but is not limited to, one or more rooms.
  • the physical space 160 can include, but is not limited to, a living room, bedroom, study, etc., or a combination of one or more of the above.
  • the physical space 160 can include, but is not limited to, a classroom, laboratory, library, etc.
  • the robot device 110 may include multiple parts.
  • the control unit 111 can serve as the control center of the robot device 110, and an application can be loaded into the control unit 111 to control the various parts of the robot device.
  • the user 120 can use the interaction unit 112 to interact with the robot device 110, for example, by inputting control commands to the robot device 110 to perform desired tasks.
  • the robot device 110 may include an arm 113 for performing actions such as grasping and releasing.
  • the arm 113 can grasp an object and move it to a desired position, and so on.
  • the robot device 110 may also include a data acquisition unit 114.
  • the data acquisition unit 114 may include various types, such as an image acquisition unit, a sound acquisition unit, etc.
  • the robot device 110 may further include a sensing unit for detecting surrounding objects, for example, detecting the distance between the robot and surrounding objects based on laser light, etc.
  • the robot device 110 may also include a drive unit 115; for example, the robot device 110 may be deployed on a movable base, and the drive unit 115 may drive the wheels of the base to move along a desired path.
  • the physical space 160 may include one or more acquisition units 130, ..., and 132.
  • one or more image acquisition devices may be deployed in the room to acquire images of the room from various angles.
  • the physical space 160 may include a control device 140, which can control one or more acquisition units 130, ..., via a network (not shown). And 132, etc.
  • control device 140 can control various electrical devices in physical space 160.
  • user 120 can instruct robot device 110 to manipulate various objects in physical space 160.
  • objects can be various items in the home environment.
  • user 120 can instruct robot device 110 to find a certain object in physical space 160; or user 120 can instruct robot device 110 to place the found object in a designated location, and so on.
  • Figure 2 illustrates a block diagram 200 for performing user tasks according to some implementations of this disclosure.
  • the robot device 110 in physical space 160 can receive a user task 210 from user 120.
  • user task 210 can instruct the robot device 110 to retrieve a first object 220.
  • the first object 220 can be an "English book,” and user 120 can say "Give me the English book” in natural language to the interaction unit 112.
  • the robot device 110 has a built-in voice recognition system.
  • the recognition module receives and parses the voice command, converting it into an executable user task 210.
  • the robot device 110 can directly give the first object 220 to the user 120, or it can identify a second object 230 (e.g., an English exercise book) that is closely related to the first object 220 mentioned in the user task 210 (e.g., an English book).
  • a second object 230 e.g., an English exercise book
  • the robot device 110 can call the machine learning model 150 to analyze the user 120's historical behavior. For example, whenever the user 120 picks up the English book, they often pick up the English exercise book to practice. Based on historical data, the model 150 predicts that the user 120 may want to obtain both the "English book” and the "English exercise book” at the same time, thus confirming the English exercise book as the second object.
  • interaction unit 112 can capture this voice command.
  • the natural language processing (NLP) algorithm built into control unit 111 begins parsing the text to understand the meaning of the command.
  • Model 150 can also access an object relation database that stores information about the relationships between various items. For example, the database might record the high frequency of co-occurrence between English books and English exercise sets, indicating a close connection between them in usage scenarios.
  • model 150 can query the database and find that the English exercise set, as an object that frequently appears alongside the English book, can be obtained simultaneously in this task, thus identifying the English exercise set as the second object 230.
  • Model 150 can also analyze the names and attributes of items to identify potential associations. For example, the name “English Exercises” itself implies that it is learning material related to "English books.” Through the keywords “English” and “exercises” in the name, the model can identify that the two objects belong to the same learning domain, thus determining that there is an association between them.
  • the control unit 111 of the robot device 110 can invoke an image acquisition device, such as a camera, from at least one of the acquisition units 114, 130, ..., and 132 to acquire a first image 240 of the physical space 160.
  • an image acquisition device such as a camera
  • the robot device 110 can scan the first physical space 160 using a camera on its head or body. (The last sentence appears to be incomplete and possibly refers to a separate point about model 150.) With this, the robot device 110 can identify specific object shapes, colors, and text from the complex first image 240, thereby accurately locating the English book and English exercise book.
  • the robot device 110 can be instructed to acquire the first object 220 and the second object 230.
  • the control unit 111 of the robot device 110 can plan a path to reduce the travel distance and time consumption.
  • the drive unit 115 of the robot device 110 then activates and moves along the planned path to the vicinity of the English book.
  • the arm 113 of the robot device 110 extends and uses the gripping device at its end to pick up the English book. Then, the robot device 110 repeats the above process, moving to the location of the English exercise book and picking it up.
  • the robot device 110 can carry the English book and English exercise book to the location of the user 120, place the first and second objects on the desk, or deliver the first object 220 and the second object 230 to the user 120. Alternatively and/or additionally, after the robot device 110 has completed its actions, it can report to the user 120 through the interaction unit 112 that the task has been completed and awaits further instructions.
  • the methods described above can be executed on any computing device with computing capabilities.
  • the methods described above can be executed using an application deployed on robot device 110.
  • an application can be deployed on control device 140 to execute the methods described above.
  • the powerful processing capabilities of model 150 can be invoked to determine a second object 230 closely associated with the first object 220, and to locate the first object 220 and the second object 230 from the first image 240.
  • robotic devices can perform user tasks in complex physical spaces.
  • the robotic device can intelligently complete the task of acquiring a first object and an associated second object.
  • the robotic device can anticipate potential subsequent user needs and provide more intelligent and human-like services. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby... Complete the expected user tasks.
  • the first image can come from at least one of the following: an acquisition unit at the robot device 110, or an acquisition unit in the first physical space.
  • FIG3 shows a block diagram 300 of the image acquisition process according to some implementations of this disclosure.
  • a first image e.g., one or more images 310 of the physical space 160 can be acquired from the acquisition unit 114 at the robot device 110. Since the robot device 110 can move freely in the physical space 160, the acquisition unit 114 can acquire images from various locations in the physical space, thereby facilitating the search for the target object.
  • a first prompt word 410 can be obtained based on user task 210 and user information 120-1 of user 120. This first prompt word 410 is used to determine the second object 230.
  • the robot device 110 can receive a first response from the machine learning model 150 to the first prompt word 410 to determine the second object 230.
  • the prompt word could be, for example, "Please determine other books closely associated with 'English book'", and the prompt word 410 can be submitted to the model 150.
  • the control unit 111 of robot device 110 can parse the user task.
  • the literal meaning of task 210 is to locate and retrieve the English book.
  • the robot device 110 can access a user information database to analyze user 120's historical behavior, interests, and usage habits. For example, the database records that user 120 has a high probability (e.g., a preset probability threshold) of simultaneously using an English book and an English exercise book when learning English.
  • the robot device 110 can obtain the first prompt word 410.
  • This prompt word 410 not only contains information from the direct instruction of user task 210 but also implicitly includes possible subsequent needs from user 120.
  • the robot device 110 sends the first prompt word 410 to the machine learning model (e.g., language learning model 420), requesting the language model 420 to make a first response based on this first prompt word 410.
  • the language model 420 analyzes and reasons, and returns a response (i.e., the first response), confirming that "English exercise set” and "English book” have a strong correlation in the current context and should be regarded as the second object 230.
  • the robot device 110 receives and parses the first response of the language model 420, confirming that "English exercise set” is the second object 230 related to "English book” in user task 210.
  • the robot device 110 can acquire and analyze prompt words and interact with the language model 420, providing the ability to handle complex tasks. Furthermore, by determining the second object 230 based on user information 120-1 and user task 210, the robot device 110's task execution capability in complex environments is enhanced, and more considerate and efficient services are provided to the user 120.
  • the first prompt word can also be obtained based on the first image.
  • the robot device 110 when the robot device 110 receives the user task 210 "Give me the English book," the robot device 110 first moves to the study and uses its acquisition unit 114, for example, through a camera or vision sensor, to capture a first image 310 of the physical space 160.
  • the robot device 110 or its connected image processing module can preprocess the first image 310, including adjusting brightness, contrast, etc., to improve image quality.
  • the object detection algorithm identifies various objects in the first image 310, including... Including, but not limited to, English books.
  • the robotic device 110 can also construct a first cue word 510, such as, "Please identify the object in the following image that is associated with the English book.”
  • the first cue word 510 not only includes a text description but also a description of the first image, which can guide the model 520 to identify other items in the image that are closely related to the English book.
  • the robot device 110 sends the first image 310 and the first cue word 510 to the model 520, requesting the language model 520 to analyze the image for a second object 230 related to an English book.
  • the model 520 uses its deep learning algorithm to analyze the first image 310 and identify items highly associated with English books, such as an English workbook. Then, the model returns a response confirming that "English workbook" is the second object 230. Finally, based on the response from the model 520, the robot device 110 confirms that the English workbook is the second object 230.
  • robot device 110 can construct a first prompt word 510 using image information, and then determine a second object associated with the first object, which helps to improve the task understanding and execution capabilities of robot device 110.
  • a second object in response to the first image indicating that the physical space includes a plurality of second objects associated with the first object, a second object can be selected from the plurality of second objects.
  • a plurality of other objects associated with the first object e.g., an English book
  • an English exercise book such as an English exercise book, an English notebook, a vocabulary book, etc.
  • machine learning models can be used to analyze the first image and user information 120-1 to predict the user 120’s possible subsequent needs, thereby making the best choice. For example, if the user is a primary school student, then English books and workbooks for primary school level will be provided to that user first. Similarly, if the user is a secondary school student, then English books and workbooks for secondary school level will be provided to that user first.
  • the robot device can not only identify multiple second objects associated with the first object in the physical space 160, but also select the second object that best meets the user's needs from the multiple second objects, which helps to improve the efficiency and accuracy of the robot device 110 in performing its tasks.
  • the robot device 110 when the robot device 110 confirms that the first object exists but the second object does not exist, the robot device 110 will directly focus on acquiring the first object without performing an additional search for the second object. This allows the robot device 110 to accurately execute the user task 210 and effectively utilize resources.
  • a message can also be provided to user 120 to inquire about the location of the second object in physical space 160.
  • the robot device 110 can be instructed to retrieve the second object based on the response.
  • robot device 110 confirms the presence of a first object (e.g., an English book) by analyzing a first image in physical space 160, but fails to identify a related second object (e.g., an English workbook) in the image.
  • Robot device 110 can further interact intelligently with user 120, supplementing the lack of environmental perception through user 120's responses.
  • user 120 receives the message
  • information about the location of the second object can be provided.
  • user 120's response to the message could be, "The English exercise book is on the table in the study.” “Up.”
  • user 120's response contains the key information needed for robot device 110 to complete the task.
  • robot device 110 can incorporate this new information into its task planning, replan its path, and find and obtain the second object, such as an English exercise book.
  • the robot device 110 can dynamically adjust its ability to perform tasks based on feedback from the user 120 during task execution.
  • user 120 can also control robot device 110 to read text aloud to assist user 120 in reading.
  • the first object is a book
  • user 120 can request robot device 110 to convert the text information in the book into speech.
  • the request from user 120 can be completed via voice command, touchscreen operation, or any other user interface.
  • robot device 110 upon receiving a request from user 120, robot device 110 begins to recognize the text content in the book. For example, robot device 110 scans the book pages using optical character recognition technology and converts the printed text into digital text format, thus forming first text data. Robot device 110 then converts the first text data into first audio data. For example, robot device 110 uses text-to-speech (TTS) technology to synthesize the text information into human-readable speech, allowing user 120 to hear the content of the book without having to read it directly.
  • TTS text-to-speech
  • the robot device 110 can complete the conversion from book text to voice output after receiving instructions from the user 120. This can provide a way for visually impaired users 120 to obtain written information, and can also allow busy users or those who prefer listening to books to access book content while doing other things, thus expanding the ways to access book content and the usage scenarios.
  • the robot device 110 when requesting the robot device 110 to convert text information in a book into speech, it can also be requested to convert text data at a specific location in the book into audio data.
  • user 120 can issue a specific request to robot device 110, and robot device 110 determines the location of the first text data in the book based on the request.
  • the first text data at the corresponding location is converted into audio data. Audio data.
  • the first text data at the corresponding location could be, for example, a specific chapter, paragraph, page, or even several lines of text.
  • User 120 can issue this request via voice command, touchscreen input, or other interactive methods, and the request may include the book's title, author's name, page number, chapter title, or keywords, etc.
  • robot device 110 converts the text data at the corresponding location into voice output. This method can enhance the application prospects of robot device 110 in fields such as reading assistance, information retrieval, and entertainment, and can provide users 120 with more convenient, efficient, and personalized reading experiences for different needs.
  • the robot device 110 can also assist the user 120 in consulting the content of the second object, thereby improving the user 120's learning efficiency. For example, the robot device 110 can identify second text data associated with the first text data in the second object. Then, the second text data is provided to the user 120.
  • robot device 110 can further play its auxiliary role by supplementing and deepening user 120's learning experience by consulting the content of a second object.
  • robot device 110 uses its information retrieval and analysis capabilities to find second text data related to the first text data from the second object.
  • robot device 110 can present the second text data to user 120, for example, through voice reading, screen display, or other suitable output methods.
  • the robotic device 110 can not only increase the depth and breadth of user 120's learning, but also improve user 120's learning efficiency. User 120 does not need to interrupt reading to search for relevant information on their own, thereby improving their focus on learning knowledge.
  • the second text data may be, for example, exercises for a portion of a chapter, which the user 120 can practice.
  • the robotic device 110 can provide a series of exercises related to the current learning content, which can help the user 120 consolidate and test their understanding and mastery of the learned knowledge.
  • robot device 110 After user 120 completes the exercises and submits the answers, robot device 110 will receive...
  • the robot device 110 analyzes the responses, such as checking the correctness, completeness, and rationality of the solution approach. It can also evaluate each answer submitted by the user 120, for example, determining whether the answer is correct or incorrect.
  • the robot device 110 can help the user 120 to instantly check their learning results by providing exercises closely related to the learning content, while the automatic grading function ensures that the user 120 can obtain timely and accurate feedback, thereby promoting the improvement of learning effectiveness and self-correction.
  • a motion model can be used to determine the specific actions to be performed by the robot device. See Figure 6 for further details, which shows a block diagram 600 illustrating the process of invoking a motion model according to some implementations of this disclosure. As shown in Figure 6, a motion model 630 can be provided, which can determine the specific actions to be performed by the robot device based on the current state and instructions of the robot device. This motion model can be a pre-trained and fine-tuned model.
  • the current state may include data from multiple aspects, such as an image of the robot device, an image of the robot device's environment, pose data of the robot arm (e.g., the positions of the robot arm's joints (POS1, ...)), and the state of the tool (e.g., a gripper, a cutting tool, etc.) fixed to the end of the robot arm.
  • POS1, (7) the positions of the robot arm's joints
  • the state of the tool e.g., a gripper, a cutting tool, etc.
  • Instructions and the current state can be input into the motion model 630, which then uses the motion model to determine the action to be performed by the robot device based on the instructions and the current state.
  • the action can represent the difference between the robot device's current pose and the next pose, and the difference between the tool's current state and the next state, etc.
  • An instruction 610 (e.g., "get an English book") can be input to the motion model 630.
  • instruction 610 can be expressed in natural language, and the instruction 610 can be determined from the response of the language model.
  • the current state of the robot device can be acquired, and the motion model 630 can determine the corresponding action 640 based on the input data. For example, the orientation, position, velocity, acceleration, etc., of each joint in the arm, and/or the wheels and/or other movable parts of the robot device at the next time point can be determined. Furthermore, it is possible to...
  • the state of the robot device at the next time point is controlled by a defined action 640.
  • a relationship can be established between the language model and the action model, and the user's initial input, expressed in natural language, can be converted into specific actions that can be performed by the robotic device. In this way, the actions of the robotic device can be precisely controlled, thereby executing the user task with higher efficiency.
  • the first object if the first object is obscured by other objects, those objects can be removed first, and then the first object can be retrieved.
  • the third object in response to determining that the first image indicates the first object is obscured by a third object in the first physical space, the third object can be moved to retrieve the first object. See Figure 7 for further details, which shows a block diagram 700 of the process of moving an object according to some implementations of this disclosure.
  • the first object 720 is an English book to be retrieved, and the third object 740 is located to the left of the first object 720 and obscures it.
  • the robot device 110 can be instructed to move the third object 740 from its current position to a position that does not obstruct the retrieval of the first object 720.
  • a target location can be determined, and the robotic device can be instructed to move the third object 740 to the target location.
  • a motion model will generate actions to control the robotic device to move the third object 740 from its current location to the target location. In this way, the robotic device can handle complex problems in complex environments, thereby performing user tasks more accurately.
  • constraints can be determined based on the pose of the third object, and the robot can be instructed to move the third object under these constraints. In this way, it can be ensured that all actions of the robot in complex environments comply with safety regulations.
  • the method of organizing the bookshelf can also be determined, and the robot device 110 can be instructed to organize the remaining objects in the bookshelf according to the above-described method.
  • the robot device After the robot device has retrieved the first object 720, new images can be acquired and new prompts can be constructed to query the language model for the next instruction. Prompts can be represented, for example, as: "Please determine the next instruction based on the following image” or "What's the next step?" And so on.
  • the language model can return "arrange the remaining books neatly," at which point, based on the instruction "arrange the remaining books neatly" and the current state of the robot, a corresponding action can be generated to instruct the robot to arrange the remaining books.
  • the robot device 110 after the robot device 110 has finished organizing the remaining books, it can be instructed to retrieve the first and second objects and proceed to the location of the user 120.
  • images can be acquired in real time, and the user's position can be located within the images.
  • corresponding instructions can be determined based on the robot device's current position (e.g., position A) and the user's position (e.g., position B). The instruction can then be expressed as: move from position A to position B.
  • the motion model 630 will generate a corresponding action that controls the robot device 110 to move from position A to position B along a determined trajectory. In this way, the robot device 110 completes the task of "give me the English book.”
  • the robot can be controlled in environments such as Chinese, English, Japanese, and French.
  • the multilingual capabilities provided by machine learning technology can be used to control the robot in application environments in different languages.
  • the robotic device can be controlled to perform other user tasks, such as finding other items in a room, placing an item in a designated location, etc.
  • users can interact with the robot device through language, actions, gestures, etc. For example, users can state the user task they wish to perform, predefine a certain action to specify the user task, and so on.
  • the robot device can automatically ask the user if they need to obtain learning materials. If a positive response is received, the robot device can retrieve the learning materials.
  • a user can interact with the robotic device via the interaction unit 112, for example, by inputting a task represented by text and/or images, and controlling the robot.
  • the device performs the task.
  • the user can specify the execution conditions of the task, such as executing the task immediately, executing the task after a predetermined time, or executing the task when predetermined conditions are determined to be met (e.g., after the user leaves school), etc.
  • the robot device can provide users with various messages. For example, assuming the robot device finds multiple types of English books, it can ask the user which type they need. Or, assuming the robot device doesn't find any English books and only finds math books, it can ask the user if they need a math book, and so on. Alternatively and/or additionally, the robot device can ask the user where the desired object can be found and then go to the user-specified location to find the desired object. Alternatively and/or additionally, if the desired object cannot be found, the robot device can ask the user if they want to purchase it, and so on.
  • various positioning algorithms can be used to determine the position of robotic devices and various objects in the physical environment.
  • a Global Positioning System GPS
  • satellite signals can be used to determine the precise position of the robotic device.
  • a communication unit can be deployed at the robotic device, and the position of the robotic device can be determined by means of signals between the communication unit and a base station and by utilizing a communication network.
  • a Wi-Fi access point can be deployed in the physical space, and the communication unit at the robotic device can interact with the Wi-Fi hotspot to determine the position via Wi-Fi signal strength and the known location of the Wi-Fi access point.
  • the communication unit at the robotic device can support Bluetooth functionality, in which case Bluetooth signals and the known locations of Bluetooth devices can be used to determine the position of nearby devices.
  • An inertial navigation system can be deployed at the robotic device, and accelerometers and gyroscopes can be used to measure and calculate the movement and orientation of the device in space, thereby determining the position of the robotic device.
  • Alternate and/or additional locations can be determined using a visual positioning system to locate the robotic device and/or individual objects.
  • a map of the physical space can be pre-acquired, and the locations of each object can be marked on this map.
  • the robotic device can use an echo detection unit to detect the distance to surrounding objects, and by combining the acquired images and the physical space map, the specific location of each object can be determined.
  • CAD computer-aided design
  • GIS geographic information systems
  • tracking units can be deployed at important objects in physical space. For example, tracking units can be added to the remote controls of home appliances (e.g., TV remote controls, air conditioner remote controls) so that the robotic device can obtain the precise location of important objects in a timely manner, and so on.
  • the robot's initial position and desired destination can be determined based on the methods described above.
  • the robot can determine a path from its initial position to its destination. For example, it can continuously acquire images of the surrounding environment and, while ensuring obstacle avoidance, continuously update the path, enabling the robot to move along the path to its destination.
  • the robotic device can perform a specified task. For example, it can acquire a specified object and move it to the appropriate location.
  • Constraints i.e., the constraints that should be followed during task execution, can be determined using a language model and/or a knowledge base. For example, an image and corresponding prompts can be acquired, and the image and prompts can be input into the language model, thereby receiving the constraints from the language model.
  • prompts can be determined as: "Based on the following image, determine the constraints that should be followed during the movement of object XXX,” or "Please determine the precautions during the movement of object XXX,” etc.
  • an object e.g., bottled water, plate, bowl, etc.
  • the object's original posture should be maintained (e.g., remaining vertical and not tilted).
  • constraints can be input into the motion model, at which point the series of actions output by the motion model will perform the corresponding tasks while ensuring the constraints are met.
  • safety during the operation of robotic devices can be ensured, thereby preventing accidental damage to an object, and so on.
  • robotic devices can perform user tasks in complex physical spaces.
  • the robotic device can intelligently complete the task of acquiring a first object and an associated second object.
  • the robotic device can anticipate potential subsequent user needs and provide more intelligent and human-like services. This approach improves the flexibility and accuracy of the robotic device in performing tasks in complex environments, thereby... Complete the expected user tasks.
  • Figure 8 illustrates a flowchart of a method 800 for performing a user task according to some implementations of this disclosure.
  • a user task is received from a user, instructing a robotic device to acquire a first object.
  • a second object associated with the first object is determined.
  • the robotic device acquires both the first and second objects.
  • determining the second object includes: obtaining a first prompt word based on the user task and the user's user information, the first prompt word being used to determine the second object; and receiving a first response from a machine learning model to the first prompt word in order to determine the second object.
  • obtaining the first prompt word further includes: obtaining the first prompt word based on the first image.
  • the method 800 further includes: in response to a first image indicating that the physical space includes a plurality of second objects associated with the first object, selecting a second object from the plurality of second objects.
  • the method 800 further includes: in response to a first image indicating that a first physical space includes a first object but does not include a second object, a robotic device acquires the first object.
  • the method 800 further includes: providing a message to a user, the message being used to inquire with the user about the location of the second object in physical space; and in response to receiving a response from the user to the message, the robotic device acquiring the second object based on the response.
  • the first object is a book
  • the method 800 further includes: receiving a request from a user to provide voice data associated with the book; the robotic device recognizing first text data in the book; and converting the first text data into... First audio data.
  • the method 800 further includes: determining the position of the first text data in the book based on a request; and converting the first text data at the position into first audio data.
  • the second object is a book
  • the method 800 further includes: determining second text data associated with the first text data in the second object; and providing the second text data to the user.
  • the method 800 further includes: in response to receiving a user's response to the second text data, determining an evaluation of the response.
  • FIG. 9 shows a block diagram of an apparatus 900 for performing a user task according to some implementations of the present disclosure.
  • the apparatus 900 includes: a receiving module 910 configured to receive a user task from a user, the user task instructing a robot device to acquire a first object; a determining module 920 configured to determine a second object associated with the first object; and an acquiring module 930 configured to, in response to a first image indicating that the first physical space in which the robot device is located includes the first object and the second object, cause the robot device to acquire the first object and the second object.
  • the determining module 920 is further configured to: obtain a first prompt word based on the user task and the user's user information, the first prompt word being used to determine a second object; and receive a first response from a machine learning model to the first prompt word in order to determine the second object.
  • the determining module 920 is further configured to: obtain a first prompt word based on the first image.
  • the acquisition module 930 is further configured to: select a second object from the plurality of second objects in response to a first image indicating that the physical space includes a plurality of second objects associated with the first object.
  • the acquisition module 930 is further configured to: In response to a first image indicating that the first physical space includes a first object but not a second object, the robotic device acquires the first object.
  • the acquisition module 930 is further configured to: provide a message to a user, the message being used to inquire of the location of the second object in physical space; and in response to receiving a response from the user to the message, cause the robotic device to acquire the second object based on the response.
  • the first object is a book
  • the receiving module 910 is further configured to: receive a request from a user to provide voice data associated with the book, enabling the robotic device to recognize first text data in the book; and convert the first text data into first audio data.
  • the receiving module 910 is further configured to: determine the position of the first text data in the book based on a request; and convert the first text data at the position into first audio data.
  • the second object is a book
  • the receiving module 910 is further configured to: determine second text data associated with the first text data in the second object; and provide the second text data to the user.
  • the receiving module 910 is further configured to: determine the evaluation of the response in response to receiving a user's response to the second text data.
  • Figure 10 shows a block diagram of a device 1000 that can implement various implementations of the present disclosure. It should be understood that the computing device 1000 shown in Figure 10 is merely exemplary and should not constitute any limitation on the functionality and scope of the implementations described herein. The computing device 1000 shown in Figure 10 can be used to implement the methods described above.
  • the computing device 1000 is in the form of a general-purpose computing device.
  • Components of the computing device 1000 may include, but are not limited to, one or more processors or processing units 1010, memory 1020, storage devices 1030, one or more communication units 1040, one or more input devices 1050, and one or more output devices 1060.
  • the processing unit 1010 may be a physical or virtual processor and can perform various processes according to programs stored in memory 1020. In a multiprocessor system, multiple processing units execute computer-executable functions in parallel. Instructions to improve the parallel processing capabilities of computing device 1000.
  • Computing device 1000 typically includes multiple computer storage media. Such media can be any available media accessible to computing device 1000, including but not limited to volatile and non-volatile media, removable and non-removable media.
  • Memory 1020 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
  • Storage device 1030 can be removable or non-removable media and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and/or data (e.g., training data for training) and can be accessed within computing device 1000.
  • the computing device 1000 may further include additional removable/non-removable, volatile/non-volatile storage media.
  • disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
  • each drive may be connected to a bus (not shown) via one or more data media interfaces.
  • the memory 1020 may include a computer program product 1025 having one or more program modules configured to perform various methods or actions of various implementations of this disclosure.
  • the communication unit 1040 enables communication with other computing devices via a communication medium. Additionally, the functionality of the components of the computing device 1000 can be implemented as a single computing cluster or multiple computing machines that can communicate via communication connections. Therefore, the computing device 1000 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
  • PCs network personal computers
  • a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
  • a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
  • a computer program product is provided that stores a computer program thereon, which, when executed by a processor, implements the methods described above.
  • These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
  • the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby enabling the instructions that execute on the computer, other programmable data processing apparatus, or other device to be implemented.
  • each block in a flowchart or block diagram may represent a module, segment, or portion of an instruction, which contains one or more executable instructions for implementing the specified logical function.
  • the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved.
  • each block in the block diagrams and/or flowcharts, and combinations of blocks in the block diagrams and/or flowcharts may be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Robotics (AREA)
  • Mechanical Engineering (AREA)
  • Manipulator (AREA)

Abstract

一种用于执行用户任务的方法、装置、设备和介质。其中,用于执行用户任务的方法包括如下步骤:接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定与第一对象相关联的第二对象;以及响应于确定机器人设备所在的物理空间的第一图像指示第一物理空间中包括第一对象和第二对象,指示机器人设备获取第一对象和第二对象。通过该方法,机器人设备可以在复杂的物理空间中执行用户任务,可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而完成预期的用户任务。

Description

用于执行用户任务的方法、装置、设备和介质 技术领域
本公开的示例性实现方式总体涉及机器人领域,特别地涉及利用机器人来执行用户任务的方法、装置、设备和计算机可读存储介质。
背景技术
机器人技术已经得到了迅速发展,并且已经被广泛地用于多个技术领域。目前已经开发出了多种专用机器人设备,例如,在工业环境中,可以使用机器人来执行加工、抓取、分类、包装等多种任务。又例如,在家居环境下,已经开发出了扫地机器人、擦玻璃机器人,等等。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。
发明内容
在本公开的第一方面,提供了一种用于执行用户任务的方法。在该方法中,接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定与第一对象相关联的第二对象;以及响应于确定机器人设备所在的物理空间的第一图像指示第一物理空间中包括第一对象和第二对象,指示机器人设备获取第一对象和第二对象。
在本公开的第二方面,提供了一种用于执行用户任务的装置。该装置包括:接收模块,被配置用于接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定模块,被配置用于确定与第一对象相关联的第二对象;以及获取模块,被配置用于响应于确定机器人设备所在的物理空间的第一图像指示第一物理空间中包括第一对象和第二对象,指示机器人设备获取第一对象和第二对象。
在本公开的第三方面,提供了一种电子设备。该电子设备包括:至少一个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令,指令在由至少一个处理单元执行时使电子设备执行根据本公开第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,计算机程序在被处理器执行时使处理器实现根据本公开第一方面的方法。
在本公开的第五方面,提供了一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据本公开第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实现方式的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
附图说明
在下文中,结合附图并参考以下详细说明,本公开各实现方式的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标注表示相同或相似的元素,其中:
图1示出了根据本公开的一个示例性实现方式的应用环境的框图;
图2示出了根据本公开的一些实现方式的用于执行用户任务的框图;
图3示出了根据本公开的一些实现方式的图像采集过程的框图;
图4示出了根据本公开的一些实现方式的调用语言模型的过程的框图;
图5示出了根据本公开的另一些实现方式的调用语言模型的过程的框图;
图6示出了根据本公开的一些实现方式的调用动作模型的过程的 框图;
图7示出了根据本公开的一些实现方式的获取对象的过程的框图;
图8示出了根据本公开的一些实现方式的用于执行用户任务的方法的流程图;
图9示出了根据本公开的一些实现方式的用于执行用户任务的装置的框图;以及
图10示出了可以实施本公开的多个实现方式的设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实现方式。虽然附图中示出了本公开的某些实现方式,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实现方式,相反,提供这些实现方式是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实现方式仅用于示例性作用,并非用于限制本公开的保护范围。
在本公开的实现方式的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实现方式”或“该实现方式”应当理解为“至少一个实现方式”。术语“一些实现方式”应当理解为“至少一些实现方式”。下文还可能包括其他明确的和隐含的定义。如本文中所使用的,术语“模型”可以表示各个数据之间的关联关系。例如,可以基于目前已知的和/或将在未来开发的多种技术方案来获取上述关联关系。
可以理解的是,本技术方案所涉及的数据(包括但不限于数据本身、数据的获取或使用)应当遵循相应法律法规及相关规定的要求。
可以理解的是,在使用本公开各实施例公开的技术方案之前,均应当根据相关法律法规通过适当的方式对本公开所涉及个人信息的类型、使用范围、使用场景等告知用户并获得用户的授权。
例如,在响应于接收到用户的主动请求时,向用户发送提示信息,以明确地提示用户,其请求执行的操作将需要获取和使用到用户的个人信息。从而,使得用户可以根据提示信息来自主地选择是否向执行本公开技术方案的操作的电子设备、应用程序、服务器或存储介质等软件或硬件提供个人信息。
作为一种可选的但非限制性的实现方式,响应于接收到用户的主动请求,向用户发送提示信息的方式,例如可以是弹出窗口的方式,弹出窗口中可以以文字的方式呈现提示信息。此外,弹出窗口中还可以承载供用户选择“同意”或“不同意”向电子设备提供个人信息的选择控件。
可以理解的是,上述通知和获取用户授权过程仅是示意性的,不对本公开的实现方式构成限定,其他满足相关法律法规的方式也可应用于本公开的实现方式中。
在此使用的术语“响应于”表示相应的事件发生或者条件得以满足的状态。将会理解,响应于该事件或者条件而被执行的后续动作的执行时机,与该事件发生或者条件成立的时间,二者之间未必是强关联的。例如,在某些情况下,后续动作可在事件发生或者条件成立时立即被执行;而在另一些情况下,后续动作可在事件发生或者条件成立后经过一段时间才被执行。
示例环境
近年来,机器人技术和机器学习技术已经被广泛应用于多个应用场景。然而,机器人通常仅能执行预先设置的固定任务,并不能按照用户需求来执行不同的用户任务。尤其是,在复杂应用环境下,机器人设备难以确定用户需求进而执行相应的任务。
目前已经开发了执行特定任务的简单机器人设备,然而,此类简单机器人设备并不能理解复杂的用户指令,也不能在复杂的物理空间中按照用户指令来执行期望的任务。此时,期望可以以有效的方式控 制机器人的操作,进而执行期望的任务。
根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图1描述根据本公开的一个示例实现方式的应用环境,图1示出了根据本公开的一个示例性实现方式的应用环境的框图100。如图1所示,机器人设备110和用户120可以位于物理空间160中,并且用户120可以控制机器人设备110来执行多种任务。物理空间160可以包括但不限于一个或者多个房间。例如,在家居环境下,物理空间160可以包括但不限于客厅、卧室、书房,等等,或者包括以上一个或者多个的组合。又例如,在教学环境中,物理空间160可以包括但不限于教室、实验室、图书室,等等。
如图1所示,机器人设备110可以包括多个部分。例如,控制单元111可以作为机器人设备110的控制中心,可以向控制单元111中加载应用程序,以便控制机器人设备中的各个部分。用户120可以使用交互单元112来与机器人设备110交互,例如,向机器人设备110输入控制指令,以便利用机器人设备110执行期望的任务。机器人设备110可以包括手臂113,用于执行抓取、释放等动作。例如,手臂113可以抓取某个对象,并且将该对象移动至期望的位置,等等。
备选地和/或附加地,机器人设备110还可以包括采集单元114。在此,采集单元114可以包括多种类型,例如,图像采集单元、声音采集单元,等等。备选地和/或附加地,机器人设备110可以进一步包括用于检测周围物体的感测单元,例如,可以基于激光来检测机器人与周围物体的距离,等等。机器人设备110还可以包括驱动单元115,例如,机器人设备110可以被部署在可移动的基座之上,并且驱动单元115可以驱动基座的轮子来按照期望的路径移动。
物理空间160可以包括一个或者多个采集单元130、…、以及132,例如,可以在房间中部署一个或者多个图像采集设备以便从各个角度采集房间的图像。物理空间160可以包括控制设备140,该控制设备140可以经由网络(未示出)控制一个或者多个采集单元130、…、 以及132,等等。备选地和/或附加地,在智能家居环境下,控制设备140可以控制物理空间160中的各种电器设备。
备选地和/或附加地,可以提供机器学习模型(例如,模型150)来管理物理空间160。应当理解,尽管图1示出了模型150位于物理空间160内部,备选地和/或附加地,该模型150可以位于物理空间160之外的远程设备处,并且控制设备140、机器人设备110或者其他设备可以经由网络来访问远程的模型150。
模型150可以包括一个或多个模型。如果模型150包括多个模型,这多个模型可以包括多个类型的模型。模型150例如可以至少包括语言模型(LM)和动作模型。语言模型通过从大量语料中学习,可以具备问答能力。动作模型可以控制机器人设备110来执行各种动作。模型150例如还可以包括图像识别模型、文本识别模型等等。
如图1所示,用户120可以指示机器人设备110来操作物理空间160中的各种对象。在此,对象可以是家居环境中的各种物品,例如,用户120可以指令机器人设备110来在物理空间160中寻找某个对象;又例如,用户120可以指令机器人设备110来将找到的对象放置到指定位置,等等。
执行任务的概要
为了至少部分地解决现有技术中的不足,根据本公开的一个示例性实现方式,提出了一种用于执行用户任务的方法。参见图2描述根据本公开的一个示例性实现方式的概要,该图2示出了根据本公开的一些实现方式的用于执行用户任务的框图200。
如图2所示,物理空间160(也称为,第一物理空间)中的机器人设备110可以接收来自用户120的用户任务210。此时,用户任务210可以指示机器人设备110来获取第一对象220。例如,在图2的示例中第一对象220可以为“英语书”,用户120可以对着交互单元112以自然语言说出“拿英语书给我”。机器人设备110内置的语音 识别模块接收并解析这一语音指令,将其转化为可执行的用户任务210。
机器人设备110接收到用户任务210之后,可以直接将第一对象220拿给用户120,也可以确定与用户任务210中提到的第一对象220(例如,英语书)关联较为密切的第二对象230(例如,英语习题集)。为了更智能地理解任务,机器人设备110可以调用机器学习模型150来分析用户120的历史行为,例如,每当用户120拿起英语书时,往往随后会拿起英语习题集进行练习。模型150基于历史数据预测用户120可能希望同时获取“英语书”和“英语习题集”,从而确认英语习题集为第二对象。
备选地和/或附加地,当用户120向机器人设备110发出“获取英语书”的指令时,交互单元112可以捕捉到这一语音命令。控制单元111内置的自然语言处理(NLP)算法开始解析文本,理解指令的含义。模型150还可以访问对象关系数据库,该对象关系数据库可以存储各种物品之间的关联性信息。例如,数据库中可能记录了英语书和英语习题集之间的高频同时出现,表明它们在使用场景上的紧密联系。当接收到“英语书”的指令时,模型150可以查询数据库,发现英语习题集作为与英语书经常一同出现的对象,可以在本次任务中同时获取,从而可以确定英语习题集作为第二对象230。
备选地和/或附加地,模型150还可以分析物品的名称和属性,以确定潜在的关联。例如,“英语习题集”这一名称本身就暗示了它是与“英语书”相关的学习材料。通过名称中的关键词“英语”和“习题”,模型可以识别出这两个对象都属于同一学习领域,从而确定它们之间具有关联。
接下来,机器人设备110的控制单元111可以调用采集单元114、130、…、以及132中的至少任一项中的图像采集设备,如摄像头,来采集物理空间160的第一图像240。例如,机器人设备110可以通过其头部或身体上的摄像头扫描第一物理空间160。在模型150的支 持下,机器人设备110可以从复杂的第一图像240中识别出特定的物体形状、颜色和文字,从而精确定位英语书和英语习题集的位置。
响应于确定机器人设备110所在的物理空间的第一图像240指示第一物理空间160中包括第一对象220和第二对象230,可以指示机器人设备110获取第一对象220和第二对象230。例如,机器人设备110在确定了第一图像240中存在英语书和英语习题集之后,机器人设备110的控制单元111可以规划一条行进路径,以减小移动距离和时间消耗。机器人设备110的驱动单元115随后启动,沿着规划的路径移动至英语书附近。到达目的地后,机器人设备110的手臂113伸展,利用其末端的抓取装置抓起英语书。接着,机器人设备110重复上述过程,前往英语习题集的位置并将其抓起。
最后,机器人设备110可以携带英语书和英语习题集,移动至用户120所在的位置,可以将第一对象和第二对象摆放在书桌上,也可以将第一对象220和第二对象230交付至用户120。备选地和/或附加地,在机器人设备110的动作执行完成之后,可以通过交互单元112向用户120报告任务已完成,等待下一步指示。
根据本公开的一些实现方式,可以在具有计算能力的任何计算设备处执行上文描述的方法。例如,可以利用部署在机器人设备110处的应用程序来执行上文描述的方法。备选地和/或附加地,可以在控制设备140处部署应用程序,以便执行上述方法。具体地,可以调用模型150的强大处理能力来确定与第一对象220关联紧密的第二对象230,并从第一图像240中找到第一对象220和第二对象230。
利用本公开的示例性实现方式,机器人设备可以在复杂的物理空间中执行用户任务。以此方式,机器人设备通过接收用户任务、分析任务、定位对象、规划路径和执行动作,可以智能地完成获取第一对象以及获取相关联的第二对象的任务。机器人设备可以提前预测用户后续可能的需求,可以提供更加智能化和人性化的服务。以此方式,可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而 完成预期的用户任务。
执行任务的详细过程
已经描述了根据本公开的一些实现方式的概要,在下文中,将描述有关执行用户任务的更多细节。为了便于描述,在下文中仅以控制机器人设备110获取英语书和英语习题集作为示例,来描述执行用户任务的更多细节。
根据本公开的一些实现方式,第一图像可以来自以下至少任一项:机器人设备110处的采集单元、第一物理空间中的采集单元。参见图3描述图像采集的更多细节,该图3示出了根据本公开的一些实现方式的图像采集过程的框图300。如图3所示,可以从机器人设备110处的采集单元114获取物理空间160的第一图像(例如,一个或者多个图像310)。由于机器人设备110可以在物理空间160中自由移动,因而采集单元114可以采集物理空间中的各个位置的图像,进而便于寻找目标对象。
备选地和/或附加地,可以从采集单元130、…、以及132来获取物理空间160的第一图像。在此,采集单元130、…、以及132可以被预先部署在物理空间160内的指定位置,例如,书房内,等等。以此方式,可以从多个角度采集更为丰富的数据。
根据本公开的一些实现方式,在确定与第一对象220具有关联的第二对象230时,可以根据用户任务210和用户120的用户信息120-1来获取第一提示词410,该第一提示词410用于确定第二对象230。接下来,机器人设备110可以接收机器学习模型150针对第一提示词410的第一应答,以便确定第二对象230。提示词例如可以表示为:“请确定与‘英语书’关联紧密的其他书籍”,并且向模型150提交提示词410。
如图4所示,当用户120通过交互单元112发出“拿英语书给我”的用户任务210时,机器人设备110的控制单元111可以解析用户任 务210的字面意义,即定位并获取英语书。同时,机器人设备110还可以调用用户信息数据库,分析用户120的历史行为、兴趣和使用习惯。例如,数据库中记录了用户120在学习英语时,有较高的概率(例如,预设概率阈值)会同时使用英语书和英语习题集。基于用户任务210(例如,拿英语书给我)和用户信息120-1(例如,七年级-爱丽丝),机器人设备110可以获取第一提示词410。“英语书”作为核心词汇,加上用户行为分析得出的关联性高的“英语习题集”,从而构成了第一提示词410。该提示词410不仅包含了用户任务210直接指令的信息,还隐含了用户120可能的后续需求。
接下来,机器人设备110将第一提示词410发送至机器学习模型(例如,语言学习模型420),请求语言模型420基于此第一提示词410做出第一应答。语言模型420在接收到第一提示词410后,通过分析和推理,返回一个应答(也即第一应答),确认“英语习题集”与“英语书”在当前情境下具有强关联性,应视为第二对象230。机器人设备110接收并解析语言模型420的第一应答,确认“英语习题集”为与用户任务210中“英语书”相关的第二对象230。
利用本公开的一些实现方式,机器人设备110可以通过获取和分析提示词,并与语言模型420交互,提供了处理复杂任务的能力。进一步基于用户信息120-1和用户任务210确定第二对象230,增强了机器人设备110在复杂环境中的任务执行能力,也为用户120提供了更加贴心和高效的服务。
根据本公开的一些实现方式,还可以基于第一图像获取第一提示词。如图5所示,当机器人设备110接收到用户任务210“拿英语书给我”时,机器人设备110首先移动到书房,并使用其采集单元114,例如通过摄像头或视觉传感器捕捉物理空间160的第一图像310。备选地和/或附加地,机器人设备110或与其相连的图像处理模块可以对第一图像310进行预处理,包括调整亮度、对比度等,以提高图像质量。接下来,通过目标检测算法识别第一图像310中的各个物体,包 括但不限于英语书。
在确定英语书的位置后,机器人设备110还可以构建第一提示词510,例如,“请确定如下图像中的与英语书相关联的对象”。在此处,第一提示词510不仅包含文本描述,还附带第一图像的描述,可以引导模型520确认图像中与英语书紧密相关的其他物品。
接下来,机器人设备110将第一图像310和第一提示词510发送给模型520,请求语言模型520分析图像中与英语书相关的第二对象230。模型520接收到第一提示词510后,运用其深度学习算法分析第一图像310,识别与英语书关联度高的物品,例如英语习题集。然后,模型返回一个应答,确认“英语习题集”是第二对象230。最后,机器人设备110基于模型520的应答,确认英语习题集为第二对象230。
利用本公开的一些实现方式,机器人设备110可以利用图像信息构建第一提示词510,进而确定与第一对象关联的第二对象,有助于提高机器人设备110的任务理解能力和执行能力。
根据本公开的一些实现方式,响应于第一图像指示物理空间包括与第一对象相关联的多个第二对象,可以从多个第二对象中选择第二对象。例如,在第一图像中可能存在多个与第一对象(例如,英语书)相关联的其他对象,例如,英语习题集、英语笔记册、单词本等。
在确定存在多个第二对象后,机器人设备110需要进一步判断哪些第二对象是最相关的或用户120最有可能需要的。在一些实现中,机器人设备110可以基于用户兴趣来确定具体的第二对象,例如用户的历史行为表明在获取英语书时倾向于同时使用英语习题集,那么机器人设备110将优先选择英语习题集作为第二对象。在另一些实现中,可以通过分析第一图像中各个对象的相对位置和状态,机器人设备110可以推断出哪些第二对象与英语书最有可能一起被使用,例如,与第一对象位置最近的第二对象可能曾经被一起使用。
备选地和/或附加地,还可以利用机器学习模型分析第一图像和用户信息120-1,预测用户120可能的后续需求,从而做出最佳选择。 例如,如果用户为小学生,则优先向该用户提供小学阶段的英语书和习题集。又例如,如果用户为中学生,则优先向该用户提供中学阶段的英语书和习题集。
利用本公开的一些实现方式,机器人设备不仅可以识别物理空间160中与第一对象相关联的多个第二对象,还可以从多个第二对象中选出最符合用户需求的第二对象,有助于提高机器人设备110任务执行的效率和准确性。
根据本公开的一些实现方式,响应于确定物理空间160的第一图像指示第一物理空间中包括第一对象但不包括第二对象,此时可以指示机器人设备获取第一对象。例如,通过对第一图像的分析,机器人设备110可以识别出物理空间160中存在第一对象(例如,英语书),但并未检测到第二对象(例如,英语习题集)的存在。此时,可以指示机器人设备110去执行最初的任务,即只获取第一对象(英语书)。
以此方式,当机器人设备110确认第一对象存在而第二对象不存在时,机器人设备110将直接聚焦于第一对象的获取,而不进行额外的第二对象搜索,可以使机器人设备110对用户任务210精准执行以及有效利用资源。
根据本公开的一些实现方式,在确定物理空间160的第一图像指示第一物理空间中包括第一对象但不包括第二对象时,还可以向用户120提供消息,该消息用于向用户120询问第二对象在物理空间160中的位置。接下来,响应于接收到用户120针对消息的应答,可以基于应答来指示机器人设备110获取第二对象。
例如,机器人设备110通过分析物理空间160中的第一图像,确认了第一对象(例如,英语书)的存在,但未能在图像中识别到与之相关的第二对象(例如,英语习题集)。机器人设备110还可以进一步与用户120智能交互,通过用户120的应答来补充环境感知的不足。
接下来,在用户120收到消息后,可以提供关于第二对象位置的信息,例如,用户120针对消息的应答为“英语习题集在书房的桌子 上”。在这里,用户120的应答包含了机器人设备110完成任务所需的关键信息。机器人设备110接收并解析用户120的应答后,可以将这一新信息融入其任务规划中,重新规划路径,来寻找并获取第二对象,例如英语习题集。
利用本公开的一些实现方式,机器人设备110可以在任务执行中基于用户120的反馈进行动态调整其执行任务的能力。
根据本公开的一些实现方式,用户120还可以控制机器人设备110朗读课文,来帮助用户120阅读。例如,在第一对象为书籍的情况下,用户120可以向机器人设备110发出的请求,要求机器人设备110将书籍中的文本信息转换为语音形式。在一些实现中,用户120发出的请求可以通过语音命令、触摸屏操作或任何其他用户界面完成。
接下来,在接收到用户120的请求后,机器人设备110开始识别书籍中的文本内容。例如,机器人设备110通过光学字符识别技术扫描书籍页面并将印刷文本转换为数字文本格式,从而形成第一文本数据。机器人设备110接下来第一文本数据转换为第一音频数据。例如,机器人设备110通过文本转语音(TTS)技术可以将文本信息合成人类可识别的语音,从而使用户120可以听到书籍的内容,而不需要直接阅读书籍。
利用本公开的一些实现方式,机器人设备110可以在接收到用户120的指令后,完成从书籍文本到语音输出的转换,可以为视力受损的用户120提供获取书面信息的途径,也可以使忙碌或偏爱听书的用户在做其他事情的同时获取书籍内容,扩展了书籍内容的访问方式和使用场景。
根据本公开的一些实现方式,在要求机器人设备110将书籍中的文本信息转换为语音形式时,还可以要求机器人设备110将书籍中特定位置处文本数据的转换为音频数据。例如,用户120可以向机器人设备110发出具体的请求,机器人设备110基于请求确定第一文本数据在书籍中的位置。接下来,将对应位置处的第一文本数据转换为第 一音频数据。在此处,对应位置处的第一文本数据例如可以是特定的章节、段落、页面甚至是几行字。用户120可以通过语音指令、触摸屏输入或其他交互方式发出该请求,并且在该请求中可以包括书籍的标题、作者名、页码、章节标题或关键词等。
接下来,机器人设备110接收到用户120的请求后,将对应位置的文本数据转换为语音输出。以此方式,可以提高机器人设备110在阅读辅助、信息获取和娱乐等领域的应用前景,并且可以为不同需求的用户120提供更加便捷、高效和个性化的阅读体验。
根据本公开的一些实现方式,机器人设备110还可以协助用户120查阅第二对象中的内容,以提高用户120的学习效率。例如,机器人设备110可以在第二对象中确定与第一文本数据相关联的第二文本数据。接下来,向用户120提供第二文本数据。
作为示例,在机器人设备110协助用户120阅读第一文本数据的过程中,机器人设备110还可以进一步发挥其辅助作用,通过查阅第二对象中的内容来补充和深化用户120的学习体验。例如,机器人设备110利用其信息检索和分析能力,从第二对象中找出与第一文本数据相关联的第二文本数据。接下来,机器人设备110可以将第二文本数据呈现给用户120,例如通过语音朗读、屏幕显示或其他适合的输出方式。
利用本公开的一些实现方式,机器人设备110不仅可以增加用户120学习的深度和广度,还可以提高用户120的学习效率。用户120不需要中断阅读去自行查找相关资料,从而提高学习知识的专注度。
根据本公开的一些实现方式,第二文本数据例如可以是部分章节的习题,用户120可以针对这部分习题进行练习。例如,在用户120学习某一主题或章节时,机器人设备110可以提供与当前学习内容相关联的一系列习题,习题可以帮助用户120巩固和检验对所学知识的理解和掌握程度。
当用户120完成习题并提交答案后,机器人设备110会对接收到 的应答进行分析,例如检查答案的正确性、完整性以及解答思路的合理性。机器人设备110可以对用户120提交的每一道习题答案进行评判,例如,判断答案的对错,等等。
利用本公开的一些实现方式,通过提供与学习内容紧密相关的习题练习,机器人设备110可以帮助用户120即时检验学习成果,而自动批改功能则保证了用户120可以获得及时、准确的反馈,从而促进学习效果的提升和自我修正。
根据本公开的一些实现方式,可以利用动作模型来确定机器人设备执行的具体动作。参见图6描述更多细节,该图6示出了根据本公开的一些实现方式的调用动作模型的过程的框图600。如图6所示,可以提供动作模型630,该动作模型630可以基于机器人设备的当前状态和指令来确定将要由机器人设备执行的具体动作,该动作模型可以是经过预训练和微调的模型。
应当理解,在此当前状态可以包括多个方面的数据,例如,机器人设备的图像、机器人设备的环境图像、机器人手臂的姿态数据(如,机器人手臂的各个关节的位置(POS1,…))、以及被固定在机器人手臂的末端的工具(例如,夹具,刀具,等等)的状态。例如,可以使用0来表示夹具的关闭状态,并且使用1来表示夹具的开放状态。可以将指令和当前状态输入至动作模型630,进而利用动作模型,基于指令和当前状态来确定由机器人设备将要执行的动作。在此,动作可以表示机器人设备的当前姿态与下一姿态之间的差异,以及工具的当前状态与下一状态之间的差异,等等。
可以向动作模型630输入指令610(例如,“获取英语书”),在此,指令610可以是以自然语言表示的,并且可以从语言模型的应答中确定该指令610。进一步,可以获取机器人设备的当前状态,动作模型630可以基于输入的数据来确定相应的动作640。例如,可以确定手臂中的各个关节、和/或机器人设备的轮子和/或其他可移动装置在下一时间点的朝向、位置、速度、加速度,等等。进一步,可以 利用确定的动作640来控制机器人设备在下一时间点的状态。
利用本公开的一些实现方式,可以在语言模型和动作模型之间建立关联关系,并且将用户最初输入的以自然语言表达的用户任务转换为由机器人设备可执行的具体动作。以此方式,可以精确地控制机器人设备的动作,进而以更高的效率执行用户任务。
根据本公开的一些实现方式,如果第一对象被其他物品遮挡,则可以首先移除该物品,继而取出第一对象。具体地,响应于确定第一图像表示第一对象被第一物理空间中的第三对象所遮挡,可以移动第三对象以便获取第一对象。参见图7描述更多细节,该图7示出了根据本公开的一些实现方式的移动对象的过程的框图700。如图7所示,在图像710中第一对象720为待取回的英语书,第三对象740位于第一对象720左侧并且遮挡第一对象720。此时,可以指示机器人设备110来将第三对象740从当前位置移动至不妨碍取回第一对象720的位置。
根据本公开的一些实现方式,可以确定目标位置,并且指示机器人设备来将第三对象740移动至目标位置。此时,运动模型将会生成动作,以便控制机器人设备来将第三对象740从当前位置移动至目标位置。以此方式,可以支持机器人设备在复杂环境中处理复杂问题,从而以更为准确的方式执行用户任务。
根据本公开的一些实现方式,在移动第三对象的过程中,可以基于第三对象的姿态来确定在移动第三对象期间的约束条件,并且指示机器人设备在约束条件下移动第三对象。以此方式,可以确保机器人设备在复杂环境中的各个动作均符合安全规范。
根据本公开的一些实现方式,还可以确定整理书柜的方式,并且可以指示机器人设备110按照上述整理方式整理书柜内剩余的对象。具体地,在机器人设备已经取回第一对象720之后,可以采集新的图像并且构造新的提示词,以便询问语言模型下一指令。提示词例如可以表示为:“请基于如下图像确定下一指令”、“下一步怎么办”, 等等。语言模型可以返回“将其余书籍摆放整齐”,此时可以基于指令“将其余书籍摆放整齐”和机器人设备的当前状态,生成相应的动作以便指示机器人设备整理其余的书籍。
根据本公开的一些实现方式,在机器人设备110将其余书籍整理完毕之后,可以指示机器人设备110获取第一对象和第二对象,并且去往用户120所在的位置。具体地,可以实时采集图像并且在图像中定位用户位置。进一步,可以基于机器人设备的当前位置(例如,位置A)和用户位置(例如,位置B)确定相应的指令,进而确定相应的指令。此时,指令可以表示为:从位置A移动至位置B。此时,动作模型630将会生成相应的动作,该动作可以控制机器人设备110按照确定的轨迹来从位置A移动至位置B。以此方式,机器人设备110完成“拿英语书给我”的任务。
应当理解,尽管上文以中文语言环境为示例描述根据本公开的一个示例实现方式。备选地和/或附加地,可以在多种语言环境中执行根据本公开的一个示例实现方式的技术方案。例如,可以在中文、英文、日文、法文等环境中控制机器人。具体地,可以基于机器学习技术所提供的多语言能力,来在不同语言的应用环境中控制机器人。进一步,尽管上文以获取英语书和英语习题集作为示例描述了利用机器人设备执行用户任务的过程,备选地和/或附加地,可以控制机器人设备来执行其他用户任务,例如,在房间中寻找其他物品,将某个物品放置到指定位置,等等。
根据本公开的一些实现方式,用户可以经由语言、动作、手势等来与机器人设备交互。例如,用户可以说出期望执行的用户任务,预先定义某个动作来指定用户任务,等等。当从采集到的图像序列中识别出该动作时,机器人设备可以自动询问用户是否需要获取学习资料,在获得肯定答复的情况下,机器人设备可以取回学习资料。
备选地和/或附加地,用户可以经由交互单元112来与机器人设备交互,例如,用户输入以文字和/或图像表示的任务,并且控制机器人 设备执行该任务。备选地和/或附加地,用户可以指定任务的执行条件,例如,立即执行任务,在预定时间之后执行任务,或者在确定满足预定条件(例如,在用户放学之后)时执行任务,等等。
根据本公开的一些实现方式,机器人设备可以向用户提供多种消息,例如,假设机器人设备找到多个种类的英语书,可以询问用户需要哪个种类。又例如,假设机器人设备没有找到英语书,并且只找到数学书,机器人设备可以询问用户是否需要数学书,等等。备选地和/或附加地,机器人设备可以询问用户在哪里可以找到期望的对象,并且前往用户指定的位置来寻找期望的对象。备选地和/或附加地,如果不能找到期望的对象,机器人设备可以询问用户是否需要购买,等等。
根据本公开的一些实现方式,可以利用多种定位算法来确定机器人设备、以及各个对象在物理环境中的位置。例如,可以在机器人设备处部署全球定位系统(GPS),并且使用卫星信号来确定机器人设备的精确位置。备选地和/或附加地,可以在机器人设备处部署通信单元,借助于该通信单元与基站之间的信号,并利用通信网络来确定机器人设备的位置。备选地和/或附加地,在物理空间中可以部署Wi-Fi接入点,机器人设备处的通信单元可以与Wi-Fi热点交互,以便经由Wi-Fi信号强度和已知Wi-Fi接入点的位置来确定位置。备选地和/或附加地,机器人设备处的通信单元可以支持蓝牙功能,此时可以使用蓝牙信号和已知的蓝牙设备位置来确定附近设备的位置。可以在机器人设备处部署惯性导航系统,并且使用加速度计和陀螺仪来测量和计算设备在空间中的移动和方向,进而确定机器人设备的位置。
备选地和/或附加地,可以使用视觉定位系统来确定机器人设备和/或各个对象的位置。可以预先获取物理空间的地图,并且在该地图中标注各个对象的位置。机器人设备可以利用回波检测单元来检测与周围对象之间的距离,并且结合采集的图像和物理空间的地图,来确定各个对象的具体位置。具体地,可以使用计算机辅助设计(CAD)和地理信息系统(GIS),并且利用定位算法来确定位置。备选地和/或 附加地,可以在物理空间中的重要对象处部署跟踪单元,例如,可以在家用电器的遥控器(例如,电视遥控器、空调遥控器)处添加跟踪单元,以便机器人设备可以及时获取重要对象的精确位置,等等。
根据本公开的一些实现方式,可以基于上文描述的方法来确定机器人设备自身的原始位置和期望去往的目的地位置。机器人设备可以确定从原始位置到目的地位置的路径。例如,可以不断获取周围的环境图像,并且在确保躲避障碍的情况下,不断更新该路径,并且使得机器人设备沿着该路径移动至目的地位置。
根据本公开的一些实现方式,在到达目的地位置之后,机器人设备可以执行指定的任务。例如,可以获取指定的对象并且将该对象移动至相应位置。可以利用语言模型和/或知识库来确定约束条件,也即在执行任务期间应当遵循的约束条件。例如,可以获取图像和相应提示词,并且向语言模型输入该图像和提示词,进而从语言模型接收约束条件。例如,可以确定提示词:“请基于如下图像,确定在移动XXX对象期间应当遵循的约束条件”,或者“请确定移动XXX对象期间的注意事项”,等等。
此时,可以确定在移动对象(例如,瓶装水、盘子、碗等)期间,应当保持对象的原始姿态(例如,保持竖直方向,不会被倾斜)。进一步,可以向动作模型输入约束条件,此时,动作模型输出的一系列动作将会在确保约束条件的情况下,执行相应的任务。利用本公开的一些实现方式,可以确保机器人设备操作期间的安全性,从而避免造成意外损坏某个对象,等等。
利用本公开的示例性实现方式,机器人设备可以在复杂的物理空间中执行用户任务。以此方式,机器人设备通过接收用户任务、分析任务、定位对象、规划路径和执行动作,可以智能地完成获取第一对象以及获取相关联的第二对象的任务。机器人设备可以提前预测用户后续可能的需求,可以提供更加智能化和人性化的服务。以此方式,可以提高机器人设备在复杂环境下执行任务的灵活度和精确度,进而 完成预期的用户任务。
示例过程
图8示出了根据本公开的一些实现方式的用于执行用户任务的方法800的流程图。在框810处,接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象。在框820处,确定与第一对象相关联的第二对象。在框830处,响应于确定机器人设备所在的物理空间的第一图像指示第一物理空间中包括第一对象和第二对象,机器人设备获取第一对象和第二对象。
根据本公开的一些实现方式,确定第二对象包括:基于用户任务和用户的用户信息,获取第一提示词,第一提示词用于确定第二对象;以及接收机器学习模型针对第一提示词的第一应答,以便确定第二对象。
根据本公开的一些实现方式,获取第一提示词进一步包括:基于第一图像获取第一提示词。
根据本公开的一些实现方式,该方法800进一步包括:响应于第一图像指示物理空间包括与第一对象相关联的多个第二对象,从多个第二对象中选择第二对象。
根据本公开的一些实现方式,该方法800进一步包括:响应于确定物理空间的第一图像指示第一物理空间中包括第一对象但不包括第二对象,机器人设备获取第一对象。
根据本公开的一些实现方式,该方法800进一步包括:向用户提供消息,消息用于向用户询问第二对象在物理空间中的位置;以及响应于接收到用户针对消息的应答,机器人设备基于应答来获取第二对象。
根据本公开的一些实现方式,第一对象为书籍,并且该方法800进一步包括:接收来自用户的提供与书籍相关联的语音数据的请求,机器人设备识别书籍中的第一文本数据;以及将第一文本数据转换为 第一音频数据。
根据本公开的一些实现方式,该方法800进一步包括:基于请求确定第一文本数据在书籍中的位置;以及将位置处的第一文本数据转换为第一音频数据。
根据本公开的一些实现方式,第二对象为书籍,并且该方法800进一步包括:在第二对象中确定与第一文本数据相关联的第二文本数据;以及向用户提供第二文本数据。
根据本公开的一些实现方式,该方法800进一步包括:响应于接收到用户针对第二文本数据的应答,确定应答的评价。
示例装置和设备
图9示出了根据本公开的一些实现方式的用于执行用户任务的装置900的框图。该装置900包括:接收模块910,被配置用于接收来自用户的用户任务,用户任务指示机器人设备来获取第一对象;确定模块920,被配置用于确定与第一对象相关联的第二对象;以及获取模块930,被配置用于响应于确定机器人设备所在的物理空间的第一图像指示第一物理空间中包括第一对象和第二对象,使得机器人设备获取第一对象和第二对象。
根据本公开的一些实现方式,确定模块920进一步被配置用于:基于用户任务和用户的用户信息,获取第一提示词,第一提示词用于确定第二对象;以及接收机器学习模型针对第一提示词的第一应答,以便确定第二对象。
根据本公开的一些实现方式,确定模块920进一步被配置用于:基于第一图像获取第一提示词。
根据本公开的一些实现方式,获取模块930进一步被配置用于:响应于第一图像指示物理空间包括与第一对象相关联的多个第二对象,从多个第二对象中选择第二对象。
根据本公开的一些实现方式,获取模块930进一步被配置用于: 响应于确定物理空间的第一图像指示第一物理空间中包括第一对象但不包括第二对象,使得机器人设备获取第一对象。
根据本公开的一些实现方式,获取模块930进一步被配置用于:向用户提供消息,消息用于向用户询问第二对象在物理空间中的位置;以及响应于接收到用户针对消息的应答,使得机器人设备基于应答来获取第二对象。
根据本公开的一些实现方式,第一对象为书籍,并且接收模块910进一步被配置用于:接收来自用户的提供与书籍相关联的语音数据的请求,使得机器人设备识别书籍中的第一文本数据;以及将第一文本数据转换为第一音频数据。
根据本公开的一些实现方式,接收模块910进一步被配置用于:基于请求确定第一文本数据在书籍中的位置;以及将位置处的第一文本数据转换为第一音频数据。
根据本公开的一些实现方式,第二对象为书籍,并且接收模块910进一步被配置用于:在第二对象中确定与第一文本数据相关联的第二文本数据;以及向用户提供第二文本数据。
根据本公开的一些实现方式,接收模块910进一步被配置用于:响应于接收到用户针对第二文本数据的应答,确定应答的评价。
图10示出了可以实施本公开的多个实现方式的设备1000的框图。应当理解,图10所示出的计算设备1000仅仅是示例性的,而不应当构成对本文所描述的实现方式的功能和范围的任何限制。图10所示出的计算设备1000可以用于实现上文描述的方法。
如图10所示,计算设备1000是通用计算设备的形式。计算设备1000的组件可以包括但不限于一个或多个处理器或处理单元1010、存储器1020、存储设备1030、一个或多个通信单元1040、一个或多个输入设备1050以及一个或多个输出设备1060。处理单元1010可以是实际或虚拟处理器并且可以根据存储器1020中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行 指令,以提高计算设备1000的并行处理能力。
计算设备1000通常包括多个计算机存储介质。这样的介质可以是计算设备1000可访问的任何可以获得的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器1020可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备1030可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以可以用于存储信息和/或数据(例如用于训练的训练数据)并且可以在计算设备1000内被访问。
计算设备1000可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图10中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器1020可以包括计算机程序产品1025,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实现方式的各种方法或动作。
通信单元1040实现通过通信介质与其他计算设备进行通信。附加地,计算设备1000的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器可以通过通信连接进行通信。因此,计算设备1000可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备1050可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备1060可以是一个或多个输出设备,例如显示器、扬声器、打印机等。计算设备1000还可以根据需要通过通信单元1040与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与计算设备1000交互的设备进 行通信,或者与使得计算设备1000与一个或多个其他计算设备通信的任何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,提供了一种计算机程序产品,其上存储有计算机程序,程序被处理器执行时实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现 流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (14)

  1. 一种用于执行用户任务的方法,包括:
    接收来自用户的用户任务,所述用户任务指示机器人设备来获取第一对象;
    确定与所述第一对象相关联的第二对象;以及
    响应于确定所述机器人设备所在的物理空间的第一图像指示所述物理空间中包括所述第一对象和所述第二对象,所述机器人设备获取所述第一对象和所述第二对象。
  2. 根据权利要求1所述的方法,其中确定所述第二对象包括:
    基于所述用户任务和所述用户的用户信息,获取第一提示词,所述第一提示词用于确定所述第二对象;以及
    接收机器学习模型针对所述第一提示词的第一应答,以便确定所述第二对象。
  3. 根据权利要求2所述的方法,其中获取所述第一提示词进一步包括:基于所述第一图像获取所述第一提示词。
  4. 根据权利要求2所述的方法,进一步包括:响应于所述第一图像指示所述物理空间包括与所述第一对象相关联的多个第二对象,从所述多个第二对象中选择所述第二对象。
  5. 根据权利要求1所述的方法,进一步包括:响应于确定所述物理空间的第一图像指示所述物理空间中包括所述第一对象但不包括所述第二对象,所述机器人设备获取所述第一对象。
  6. 根据权利要求5所述的方法,进一步包括:
    向所述用户提供消息,所述消息用于向所述用户询问所述第二对象在所述物理空间中的位置;以及
    响应于接收到所述用户针对所述消息的应答,所述机器人设备基于所述应答来获取所述第二对象。
  7. 根据权利要求1所述的方法,其中所述第一对象为书籍,并 且所述方法进一步包括:
    接收来自所述用户的提供与所述书籍相关联的语音数据的请求,所述机器人设备识别所述书籍中的第一文本数据;以及
    将所述第一文本数据转换为第一音频数据。
  8. 根据权利要求7所述的方法,进一步包括:
    基于所述请求确定所述第一文本数据在所述书籍中的位置;以及
    将所述位置处的所述第一文本数据转换为所述第一音频数据。
  9. 根据权利要求8所述的方法,其中所述第二对象为书籍,并且所述方法进一步包括:
    在所述第二对象中确定与所述第一文本数据相关联的第二文本数据;以及
    向所述用户提供所述第二文本数据。
  10. 根据权利要求9所述的方法,进一步包括:响应于接收到所述用户针对所述第二文本数据的应答,确定所述应答的评价。
  11. 一种用于执行用户任务的装置,包括:
    接收模块,被配置用于接收来自用户的用户任务,所述用户任务指示机器人设备来获取第一对象;
    确定模块,被配置用于确定与所述第一对象相关联的第二对象;以及
    获取模块,被配置用于响应于确定所述机器人设备所在的物理空间的第一图像指示所述物理空间中包括所述第一对象和所述第二对象,所述机器人设备获取所述第一对象和所述第二对象。
  12. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至10中任一项所述的方法。
  13. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序在被处理器执行时使所述处理器实现根据权利要求1至10中任一项所述的方法。
  14. 一种计算机程序产品,包括计算机程序,其中所述计算机程序在被处理器执行时实现根据权利要求1至10中任一项所述的方法。
PCT/CN2024/106250 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质 Pending WO2026016136A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
PCT/CN2024/106250 WO2026016136A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质
CN202480003519.7A CN121712619A (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/106250 WO2026016136A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Publications (1)

Publication Number Publication Date
WO2026016136A1 true WO2026016136A1 (zh) 2026-01-22

Family

ID=98436605

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/106250 Pending WO2026016136A1 (zh) 2024-07-18 2024-07-18 用于执行用户任务的方法、装置、设备和介质

Country Status (2)

Country Link
CN (1) CN121712619A (zh)
WO (1) WO2026016136A1 (zh)

Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112070586A (zh) * 2020-09-09 2020-12-11 腾讯科技(深圳)有限公司 基于语义识别的物品推荐方法、装置、计算机设备及介质
CN112750000A (zh) * 2019-10-30 2021-05-04 北京京东尚科信息技术有限公司 一种物品搭配的确定方法、装置、设备和存储介质
US20210248656A1 (en) * 2019-10-30 2021-08-12 Lululemon Athletica Canada Inc. Method and system for an interface for personalization or recommendation of products
CN113771048A (zh) * 2021-11-15 2021-12-10 季华实验室 一种机器人导购方法、装置、电子设备及存储介质
CN117216373A (zh) * 2023-03-13 2023-12-12 腾讯科技(深圳)有限公司 一种物品推荐方法、装置、电子设备及存储介质
CN117484502A (zh) * 2023-11-21 2024-02-02 北京有竹居网络技术有限公司 数据处理方法、装置、设备及计算机介质

Patent Citations (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112750000A (zh) * 2019-10-30 2021-05-04 北京京东尚科信息技术有限公司 一种物品搭配的确定方法、装置、设备和存储介质
US20210248656A1 (en) * 2019-10-30 2021-08-12 Lululemon Athletica Canada Inc. Method and system for an interface for personalization or recommendation of products
CN112070586A (zh) * 2020-09-09 2020-12-11 腾讯科技(深圳)有限公司 基于语义识别的物品推荐方法、装置、计算机设备及介质
CN113771048A (zh) * 2021-11-15 2021-12-10 季华实验室 一种机器人导购方法、装置、电子设备及存储介质
CN117216373A (zh) * 2023-03-13 2023-12-12 腾讯科技(深圳)有限公司 一种物品推荐方法、装置、电子设备及存储介质
CN117484502A (zh) * 2023-11-21 2024-02-02 北京有竹居网络技术有限公司 数据处理方法、装置、设备及计算机介质

Also Published As

Publication number Publication date
CN121712619A (zh) 2026-03-20

Similar Documents

Publication Publication Date Title
Hatori et al. Interactively picking real-world objects with unconstrained spoken language instructions
US10824310B2 (en) Augmented reality virtual personal assistant for external representation
Skubic et al. Spatial language for human-robot dialogs
Forbes et al. Robot programming by demonstration with situated spatial language understanding
US8893048B2 (en) System and method for virtual object placement
Daniele et al. Navigational instruction generation as inverse reinforcement learning with neural machine translation
Samadi et al. Using the web to interactively learn to find objects
WO2024059179A1 (en) Robot control based on natural language instructions and on descriptors of objects that are present in the environment of the robot
US20240096093A1 (en) Ai-driven augmented reality mentoring and collaboration
CN120063242A (zh) 一种机器人语义导航方法、系统、终端及存储介质
Constantin et al. Interactive multimodal robot dialog using pointing gesture recognition
Tan et al. Embodied scene description
VanderHoeven et al. Point target detection for multimodal communication
Hsiao et al. Object schemas for grounding language in a responsive robot
Rivkin et al. Cartier: Cartographic language reasoning targeted at instruction execution for robots
Kartmann et al. Interactive and incremental learning of spatial object relations from human demonstrations
Miura et al. Development of a personal service robot with user-friendly interfaces
WO2026016136A1 (zh) 用于执行用户任务的方法、装置、设备和介质
CN120448507A (zh) 基于多模态大模型的对话方法、装置、设备以及存储介质
Jain et al. Gesnavi: gesture-guided outdoor vision-and-language navigation
Moratz et al. Instruction modes for joint spatial reference between naive users and a mobile robot
Dizet et al. RoboCup@ Home Education 2020 Best Performance: RoboBreizh, a modular approach
Perera Language-Based Bidirectional Human and Robot Interaction Learning for Mobile Service Robots
Patel Human robot interaction with cloud assisted voice control and vision system
WO2026016144A1 (zh) 用于执行用户任务的方法、装置、设备和介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24840693

Country of ref document: EP

Kind code of ref document: A1