WO2025255789A1 - 请求处理方法、装置、设备和存储介质 - Google Patents

请求处理方法、装置、设备和存储介质

Info

Publication number
WO2025255789A1
WO2025255789A1 PCT/CN2024/099090 CN2024099090W WO2025255789A1 WO 2025255789 A1 WO2025255789 A1 WO 2025255789A1 CN 2024099090 W CN2024099090 W CN 2024099090W WO 2025255789 A1 WO2025255789 A1 WO 2025255789A1
Authority
WO
WIPO (PCT)
Prior art keywords
text
texts
candidate
query request
electronic device
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2024/099090
Other languages
English (en)
French (fr)
Inventor
王勇
杨晶生
龚笠
陈甜甜
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Douyin Vision Co Ltd
Original Assignee
Douyin Vision Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Douyin Vision Co Ltd filed Critical Douyin Vision Co Ltd
Priority to CN202480004209.7A priority Critical patent/CN121532760A/zh
Priority to PCT/CN2024/099090 priority patent/WO2025255789A1/zh
Publication of WO2025255789A1 publication Critical patent/WO2025255789A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying

Definitions

  • the exemplary embodiments disclosed herein relate generally to the field of computers, and more particularly to request processing methods, apparatus, devices, and computer-readable storage media.
  • Text retrieval is one of the most fundamental and important tasks in natural language processing. It has already been implemented in many scenarios, including intelligent question answering, intent recognition, semantic understanding, and semantic generation. For example, electronic devices can retrieve other texts that match the text content included in a query request.
  • a request processing method includes: acquiring a query request, the query request including first text; in response to a first length of the first text being greater than a threshold, generating a plurality of associated texts based on the first text, the associated texts having a second length less than the first length; determining at least one second text matching the query request from a set of candidate texts based on a plurality of first feature representations of the plurality of associated texts; and generating a response to the query request based on the at least one second text.
  • an apparatus for request processing includes: an acquisition module configured to acquire a query request, the query request including first text; a first generation module configured to generate a plurality of associated texts based on the first text in response to a first length of the first text being greater than a threshold, the associated texts having a second length less than the first length; a first determination module configured to determine at least one second text matching the query request from a set of candidate texts based on a plurality of first feature representations of the plurality of associated texts; and a second generation module configured to generate a response to the query request based on the at least one second text.
  • an electronic device in a third aspect of this disclosure, includes at least one... A processing unit; and at least one memory, coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. When executed by the at least one processing unit, the instructions cause the device to perform the method of the first aspect.
  • a computer-readable storage medium stores a computer program that can be executed by a processor to implement the method of the first aspect.
  • Figure 1 shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented
  • Figure 2 shows a flowchart of a request processing procedure according to some embodiments of the present disclosure
  • Figure 3 illustrates a schematic diagram of a request processing procedure according to some embodiments of the present disclosure
  • Figure 4 shows a schematic structural block diagram of a request processing apparatus according to certain embodiments of the present disclosure.
  • Figure 5 shows a block diagram of an electronic device capable of implementing several embodiments of the present disclosure.
  • the term “comprising” and similar terms should be understood as open-ended inclusion, i.e., “including but not limited to”.
  • the term “based on” should be understood as “at least partially based on”.
  • the term “one embodiment” or “the embodiment” should be understood as “at least one embodiment”.
  • the term “some embodiments” should be understood as “at least some embodiments”.
  • Other explicit and implicit definitions may also be included below.
  • the terms “first”, “second”, etc., may refer to different or the same objects. Other explicit and implicit definitions may also be included below.
  • the embodiments of this disclosure may involve user data, data acquisition, and/or use. All of these aspects comply with applicable laws, regulations, and relevant provisions. In the embodiments of this disclosure, all data collection, acquisition, processing, manipulation, forwarding, and use are conducted with the user's knowledge and confirmation. Accordingly, in implementing the embodiments of this disclosure, the type, scope of use, and usage scenarios of any data or information that may be involved should be communicated to the user and their authorization obtained in accordance with relevant laws and regulations through appropriate means. The specific methods of notification and/or authorization may vary depending on the actual situation and application scenario, and the scope of this disclosure is not limited in this respect.
  • any processing of personal information will be carried out only under the premise of legality (such as obtaining the consent of the personal information subject, or being necessary for the performance of a contract), and will only be carried out within the scope stipulated or agreed upon.
  • a user's refusal to process personal information other than that necessary for basic functions will not affect the user's use of basic functions.
  • Traditional text matching methods mainly include: supervised learning-based topic extraction matching methods, text segmentation-based matching methods, text semantic segmentation-based matching methods, and semantic depth matching methods. These traditional text retrieval methods only consider the word/semantic information of the text itself, without considering the problem that inappropriate segmentation can lead to semantic segmentation or semantic dispersion, resulting in low accuracy of text retrieval.
  • a query request can be obtained, the query request including first text; further, in response to a first length of the first text being greater than a threshold, multiple associated texts can be generated based on the first text, the second length of the associated texts being less than the first length; further, based on multiple first feature representations of the multiple associated texts, at least one second text matching the query request can be determined from a set of candidate texts; further, based on at least one second text, a response to the query request can be generated.
  • embodiments of this disclosure can generate multiple associated texts with a length shorter than the first text, and match at least one associated second text based on the multiple associated texts. This can effectively solve the problem of semantic dispersion caused by excessively long texts, and also improve the retrieval accuracy of texts with longer lengths.
  • Figure 1 illustrates a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented.
  • the example environment 100 may include an electronic device 110.
  • electronic device 110 can retrieve other matching text based on the first text in a query request, wherein the first text corresponds to a first length greater than a threshold (which may also be referred to as "long text” or “long text content” in this disclosure, for example).
  • the query request can be any appropriate request, which can be user input or automatically generated by the electronic device, and will not be elaborated here.
  • an electronic device when an electronic device receives a user's query for other novels similar to novel A, it can retrieve and recommend novel B, which is similar to novel A.
  • electronic devices can automatically recommend tools to bots based on information related to the bot when a user creates a bot on a bot creation platform.
  • This information can be long text content consisting of bot identifiers, bot descriptions, bot system prompts, etc.
  • the tool descriptions are equivalent to other text retrieved based on the bot-related information.
  • Electronic device 110 can be any type of mobile terminal, fixed terminal, or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, etc. Netbook computers, tablet computers, media computers, multimedia tablets, handheld computers, portable gaming terminals, VR/AR devices, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio/video players, digital cameras/camcorders, positioning devices, television receivers, radio receivers, e-book devices, gaming devices, or any combination thereof, including accessories and peripherals of these devices or any combination thereof. In some embodiments, electronic device 110 may also support any type of interface for the target user (such as "wearable" circuitry).
  • Electronic device 110 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks, and big data and artificial intelligence platforms.
  • Electronic device 110 may include, for example, computing systems/servers, such as mainframes, edge computing nodes, computing devices in cloud environments, and so on.
  • Figure 2 shows a flowchart of a request processing procedure 200 according to some embodiments of the present disclosure.
  • Procedure 200 may be implemented at electronic device 110.
  • Procedure 200 is described below with reference to Figure 1.
  • electronic device 110 can obtain a query request, which includes a first text.
  • the first text can be any suitable text, which may correspond to the base text used for the query.
  • the query request can be any suitable request, which may be user-inputted or automatically generated by the electronic device 110, and will not be elaborated here.
  • the process by which the electronic device 110 generates a response to the query request includes searching/matching based on the first text.
  • electronic device 110 may generate multiple associated texts based on the first text in response to the first text having a first length greater than a threshold, wherein the second length of the associated texts is less than the first length.
  • the electronic device 110 can generate first input information for a first model based on the first text.
  • the first input information includes a first prompt, which instructs the first model to generate associated text of a preset length based on the first text.
  • the preset length can be any appropriate length, and can be a predetermined numerical value or a predetermined length range.
  • the first prompt may instruct the first model to generate associated text with a length not exceeding a predetermined length threshold, or it may instruct the first model to generate multiple associated texts of length A based on the first text, and so on.
  • associated text can also be considered as "short text" or "short text content.”
  • the first prompt can also be used to indicate the number of associated texts generated by the first model, the format of the associated texts, how to generate associated texts based on the first text, etc., which will not be elaborated here.
  • the first prompt could be "Given a text, which may contain multiple intents. You need to understand these intents of the text and break it down into shorter texts that can represent these intents.”
  • the electronic device 110 can input the first input information into the first model to obtain multiple related texts output by the first model.
  • the second length of these multiple related texts is less than the first length of the first text, and the second lengths of these multiple related texts can be the same or different, depending on the requirements.
  • the first model can be any suitable model.
  • the first model can be a model with enhanced generalization ability and better handling of long-distance dependencies, such as a language model.
  • the electronic device 110 may use a first model to perform semantic understanding on the first text and generate multiple related texts based on the semantic understanding results.
  • the electronic device 110 can also utilize the world knowledge of the first model to perform semantic understanding of the first text, thereby generating multiple related texts.
  • the world knowledge can be any appropriate and general knowledge that assists the first model in understanding the semantics of the text. For example, world knowledge... World knowledge can include things like "there are seven continents in the world".
  • the electronic device 110 uses the first model to generate multiple related texts, which can effectively solve the problem of catastrophic forgetting caused by semantic dispersion due to text length process, and can improve the correlation between the generated multiple related texts and the first text, and can perform accurate semantic segmentation, achieving strong semantic generalization.
  • electronic device 110 can determine at least one second text that matches a query request from a set of candidate texts based on multiple first feature representations of multiple associated texts.
  • the length of the second text may be greater than the length of multiple associated texts, and the length of the second text may be greater than a predetermined threshold, that is, the second text may be text that matches the first text (e.g., long text content).
  • the electronic device 110 can determine at least one second text that matches the semantic information from a set of candidate texts based on the semantic information of multiple associated texts. That is, the electronic device 110 can determine at least one second text that matches the semantic information from a set of candidate texts based on the comparison result of the semantic information of multiple associated texts and the semantic information of this set of candidate texts.
  • the electronic device 110 can determine at least one second text matching the text information from a set of candidate texts based on the text information of multiple associated texts. That is, the electronic device 110 can determine at least one second text matching the text information from a set of candidate texts based on the comparison result of the text information of multiple associated texts and the text information of this set of candidate texts.
  • the text information can be the multiple associated texts themselves, or it can be other text information extracted based on the associated texts, such as keywords, etc.
  • the electronic device 110 can determine at least one second text matching the text information and/or semantic information from a set of candidate texts based on the textual and semantic information of multiple associated texts.
  • the electronic device 110 can determine at least one first candidate text matching the textual information from a set of candidate texts.
  • the electronic device 110 can determine at least one second candidate text matching the semantic information from a set of candidate texts.
  • the electronic device 110 can determine at least one second text matching the query request based on at least one first candidate text and at least one second candidate text.
  • the electronic device 110 can determine at least one second text matching the query request based on at least one first candidate text and at least one second candidate text. Both candidate texts are determined as the second text.
  • electronic device 110 can determine the common candidate text among at least one first candidate text and at least one second candidate text as the second text. For instance, a set of candidate texts includes candidate text A, candidate text B, candidate text C, and candidate text D. Electronic device 110 determines candidate text A and candidate text B that match the query request from the set of candidate texts based on semantic information. Electronic device 110 also determines candidate text A and candidate text C from the set of candidate texts based on textual information. Therefore, electronic device 110 can determine candidate text A as the second text.
  • the electronic device 110 may determine a first ranking result corresponding to at least one first candidate text and at least one second candidate text based on the relevance of at least one first candidate text and at least one second candidate text to a plurality of associated texts.
  • the electronic device 110 may also determine at least one second text that matches a query request based on the ranked plurality of candidate texts.
  • the electronic device can concatenate the feature representations corresponding to multiple associated texts to obtain a target feature representation, and determine the correlation between the association of at least one first candidate text and at least one second candidate text with the feature representations corresponding to multiple associated texts and the target feature representation.
  • the electronic device 110 may determine a second ranking result corresponding to at least one first candidate text and multiple second candidate texts based on their relevance to the first text.
  • the electronic device 110 may also determine at least one second text that matches the query request based on the ranked multiple candidate texts.
  • electronic device 110 can determine the target Euclidean distance based on the semantic feature representations corresponding to each candidate text and the semantic feature representations corresponding to the first text. Furthermore, electronic device 110 can determine the relevance between each candidate text and the first text based on the target Euclidean distance, where a smaller target Euclidean distance indicates a higher relevance.
  • the electronic device 110 can determine at least one second text that matches the query request based on attribute information of at least one first candidate text and a plurality of second candidate texts.
  • the attribute information can be any suitable information, such as Popularity, user feedback, etc. Taking the second text as a novel as an example, user feedback can include the number of times users have viewed it, the number of times users have liked it, the number of times users have saved it, and so on.
  • the electronic device 110 can acquire a set of candidate texts.
  • the electronic device 110 can input the set of candidate texts into a second model to acquire multiple descriptive texts corresponding to each candidate text in the set of candidate texts output by the second model, wherein the length of the descriptive texts is less than the length of the candidate texts.
  • the second model can be any suitable model.
  • the second model can be a model with enhanced generalization ability and good handling of long text content. For example, it can be a language model, etc.
  • the electronic device 110 may utilize a second model to perform semantic understanding on each candidate text and generate multiple descriptive texts based on the semantic understanding results. In some embodiments, the electronic device 110 may also utilize the world knowledge of the second model to perform semantic understanding on each candidate text to generate multiple descriptive texts.
  • Electronic device 110 can determine the semantic information of multiple descriptive texts corresponding to each candidate text in a set of candidate texts. Electronic device 110 can also determine the textual information of multiple descriptive texts corresponding to each candidate text in a set of candidate texts.
  • the electronic device 110 can also store the correspondence between candidate text, description text corresponding to candidate text, semantic information of description text corresponding to candidate text, and textual information of description text corresponding to candidate text in a vector database, so as to support the electronic device 110 to match or retrieve at least one second text that matches the query request based on the semantic information and/or textual information of multiple associated texts generated based on the first text in the query request.
  • the electronic device 110 may determine one or more descriptive texts that match the associated text in terms of semantic information and/or textual information. Further, the electronic device 110 may determine candidate texts corresponding to the one or more descriptive texts as second texts that match the long text.
  • electronic device 110 can generate a response to a query request based on at least one second text.
  • the electronic device 110 can be directly based on this at least one second document. This generates a response to the query request, which determines that at least one second text is a match for the first text.
  • the electronic device 110 may also use any one of the at least one second text as the target text and generate a response to the query request based on the target text.
  • the electronic device 110 may also provide a first text, a plurality of related texts, and at least one second text to a third model.
  • the electronic device 110 may obtain target text determined by the third model from at least one second text, the target text being determined based on the first text and the plurality of related texts.
  • the target text is text that the electronic device 110 further determines from at least one second text that is more relevant to the first text, based on the first text and the plurality of related texts.
  • the electronic device 110 may utilize the target text to generate a response to a query request.
  • the third model may be instructed to determine the target text from at least one second text based on the relevance of at least one second text to the first text and a plurality of associated texts.
  • the electronic device 110 may generate second input information input to the third model based on the first text, a plurality of associated texts, and at least one second text, wherein the second input information includes a second prompt item.
  • the second prompt can be used to instruct the third model how to determine the target text from at least one second text, etc., which will not be elaborated here.
  • the second prompt could be: "You are given a long text and some intentions corresponding to this long text. These intentions do not necessarily cover all the intentions corresponding to the text. In addition, you are given some candidate texts. You need to understand this text and find the text that matches this long text from these candidate texts based on this text and the multiple intentions corresponding to this text.”
  • Electronic device 110 can use this second input information to obtain the target text determined by the third model based on this second input information.
  • FIG. 3 illustrates a request processing procedure according to some embodiments of the present disclosure, and will now be described with reference to Figure 3.
  • Electronic device 110 can obtain query requests, which may include long texts exceeding a threshold in length. Electronic device 110 can utilize a language model to understand the semantic information corresponding to the long texts. The electronic device 110 generates m short texts with lengths less than or equal to a small threshold. The electronic device 110 can obtain the semantic representations (semantic information) corresponding to these m short texts. Based on the m long texts and the multiple descriptive texts corresponding to each candidate text stored in the vector database, the electronic device 110 can determine m*k first matching texts. The electronic device 110 can also determine m*k first matching texts based on the semantic representations corresponding to the m long texts and the semantic representations of the multiple descriptive texts corresponding to each candidate text stored in the vector database.
  • the multiple descriptive texts corresponding to each candidate text in the vector database can also be obtained after understanding the semantic information of each candidate text based on the language model.
  • the method of obtaining the semantic representation of the multiple descriptive texts corresponding to each candidate text can be the same as or different from the method of obtaining the semantic representation of the m short texts corresponding to the long text. This will not be elaborated here.
  • electronic device 110 can filter out a second matching texts from these (m+m)*k matching texts based on a model, where a is less than (m+m)*k. Furthermore, electronic device 110 can utilize a language model to perform further matching based on the long text, m short texts, and a second matching texts in the query request, obtaining the target matching text that matches the long text, and generating a response to the query request based on the target matching text.
  • the electronic device 110 can be pre-configured with parameter configuration information, which may include, but is not limited to, the aforementioned m, k, and a, etc.
  • embodiments of this disclosure can generate multiple associated texts with a length shorter than the first text, and retrieve at least one second text associated with the first text based on the multiple associated texts. This can effectively solve the problem of semantic dispersion caused by excessively long texts, and also improve the accuracy of retrieving texts with a large length.
  • FIG. 4 shows a schematic structural block diagram of an apparatus 400 for request processing according to certain embodiments of this disclosure.
  • Apparatus 400 may be implemented as or included in the electronic device 110 discussed above.
  • the various modules/components in apparatus 400 may be hardware, software, or fixed-line components. It can be achieved by components or any combination thereof.
  • the device 400 includes an acquisition module 410 configured to acquire a query request, the query request including first text; a first generation module 420 configured to generate multiple associated texts based on the first text in response to a first length of the first text being greater than a threshold, the associated texts having a second length less than the first length; a first determination module 430 configured to determine at least one second text matching the query request from a set of candidate texts based on multiple first feature representations of the multiple associated texts; and a second generation module 440 configured to generate a response to the query request based on at least one second text.
  • an acquisition module 410 configured to acquire a query request, the query request including first text
  • a first generation module 420 configured to generate multiple associated texts based on the first text in response to a first length of the first text being greater than a threshold, the associated texts having a second length less than the first length
  • a first determination module 430 configured to determine at least one second text matching the query request from a set of candidate texts based on multiple first feature representations of the multiple
  • the first generation module 420 is specifically configured to generate first input information for the first model based on the first text, the first input information including a first prompt item, the first prompt item being used to instruct the first model to generate associated text of a preset length based on the first text; and to obtain multiple associated texts output by the first model.
  • the first determining module 420 is specifically configured to determine at least one second text that matches the text information and/or semantic information from a set of candidate texts based on the semantic information and/or text information of a plurality of associated texts.
  • the first determining module 420 is specifically configured to determine at least one first candidate text that matches text information from a set of candidate texts; determine at least one second candidate text that matches semantic information from a set of candidate texts; and determine at least one second text that matches a query request based on at least one first candidate text and at least one second candidate text.
  • the apparatus 400 further includes a semantic determination module configured to: acquire a set of candidate texts; acquire multiple descriptive texts corresponding to each candidate text in the set of candidate texts output by the second model by inputting the set of candidate texts into the second model, wherein the length of the descriptive texts is less than that of the candidate texts; and determine the semantic information of the corresponding candidate texts based on the multiple descriptive texts.
  • a semantic determination module configured to: acquire a set of candidate texts; acquire multiple descriptive texts corresponding to each candidate text in the set of candidate texts output by the second model by inputting the set of candidate texts into the second model, wherein the length of the descriptive texts is less than that of the candidate texts; and determine the semantic information of the corresponding candidate texts based on the multiple descriptive texts.
  • the first determining module 420 is specifically configured to sort at least one first candidate text and at least one second candidate text based on the relevance of each candidate text to multiple associated texts and/or first text; and to determine at least one second text that matches the query request based on the sorted candidate texts.
  • the first determining module 420 is specifically configured to determine at least one second text that matches the query request from at least one first candidate text and at least one second candidate text based on the attribute information of each candidate text.
  • the second generation module 440 is specifically configured to provide a first text, multiple associated texts, and at least one second text to a third model;
  • target text determined by a third model from at least one second text, the target text being determined based on a first text and multiple related texts; and use the target text to generate a response to the query request.
  • the third model is instructed to determine the target text from at least one second text based on the relevance of at least one second text to the first text and a plurality of associated texts.
  • the units included in device 400 can be implemented in various ways, including software, hardware, firmware, or any combination thereof.
  • one or more units may be implemented using software and/or firmware, such as machine-executable instructions stored on a storage medium.
  • some or all of the units in device 400 may be implemented at least partially by one or more hardware logic components.
  • exemplary types of hardware logic components include field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-chips (SoCs), complex programmable logic devices (CPLDs), and so on.
  • Figure 5 shows a block diagram of an electronic device 500 in which one or more embodiments of the present disclosure may be implemented. It should be understood that the electronic device 500 shown in Figure 5 is merely exemplary and should not constitute any limitation on the functionality and scope of the embodiments described herein. The electronic device 500 shown in Figure 5 can be used to implement the electronic device 110 shown in Figure 1.
  • the electronic device 500 is in the form of a general-purpose electronic device.
  • Components of the electronic device 500 may include, but are not limited to, one or more processors or processing units 510, memory 520, storage devices 530, one or more communication units 540, one or more input devices 550, and one or more output devices 560.
  • the processing unit 510 may be a physical or virtual processor and is capable of performing various processes according to programs stored in memory 520. In a multiprocessor system, multiple processing units execute computer-executable instructions in parallel to improve power efficiency.
  • the parallel processing capability of the sub-device is 500.
  • Electronic device 500 typically includes multiple computer storage media. Such media can be any accessible media that is accessible to electronic device 500, including but not limited to volatile and non-volatile media, removable and non-removable media.
  • Memory 520 can be volatile memory (e.g., registers, cache, random access memory (RAM)), non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof.
  • Storage device 530 can be a removable or non-removable medium and can include machine-readable media, such as flash drives, disks, or any other media that can be used to store information and/or data (e.g., training data for training) and can be accessed within electronic device 500.
  • Electronic device 500 may further include additional removable/non-removable, volatile/non-volatile storage media.
  • disk drives for reading from or writing to removable, non-volatile disks (e.g., "floppy disks") and optical disk drives for reading from or writing to removable, non-volatile optical disks may be provided.
  • each drive may be connected to a bus (not shown) via one or more data media interfaces.
  • Memory 520 may include computer program product 525 having one or more program modules configured to perform various methods or actions of various embodiments of the present disclosure.
  • Communication unit 540 enables communication with other electronic devices via a communication medium. Additionally, the functionality of components of electronic device 500 can be implemented using a single computing cluster or multiple computing machines capable of communicating via communication connections. Therefore, electronic device 500 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network node.
  • PCs network personal computers
  • Input device 550 can be one or more input devices, such as a mouse, keyboard, trackball, etc.
  • Output device 560 can be one or more output devices, such as a monitor, speaker, printer, etc.
  • Electronic device 500 can also communicate with one or more external devices (not shown) via communication unit 540 as needed.
  • External devices include storage devices, display devices, etc., and can communicate with one or more devices that enable user interaction with electronic device 500, or with any device that enables electronic device 500 to communicate with one or more other electronic devices. This allows communication with any device (e.g., network card, modem, etc.). Such communication can be performed via an input/output (I/O) interface (not shown).
  • I/O input/output
  • a computer-readable storage medium that stores computer-executable instructions thereon, wherein the computer-executable instructions are executed by a processor to implement the methods described above.
  • a computer program product is also provided, which is tangibly stored on a non-transitory computer-readable medium and includes computer-executable instructions, which are executed by a processor to implement the methods described above.
  • These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when executed by the processing unit of the computer or other programmable data processing apparatus, they create means for implementing the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, programmable data processing apparatus, and/or other device to operate in a particular manner.
  • the computer-readable medium storing the instructions comprises an article of manufacture that includes instructions for implementing aspects of the functions/actions specified in one or more blocks of the flowchart and/or block diagram.
  • Computer-readable program instructions can be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions that execute on the computer, other programmable data processing apparatus, or other device to perform the functions/actions specified in one or more boxes of a flowchart and/or block diagram.
  • each block in the flowchart or block diagram may represent a module, program segment, or part of an instruction.
  • a block, module, program segment, or part of an instruction contains one or more executable instructions for implementing a specified logical function.
  • the functions marked in the blocks may occur in a different order than those shown in the figures. For example, two consecutive blocks may actually be executed substantially in parallel, or sometimes in reverse order, depending on the functions involved.
  • each block in a block diagram and/or flowchart, and combinations of blocks in block diagrams and/or flowcharts can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

根据本公开的实施例,提供了请求处理方法、装置、设备和存储介质。该方法包括:获取查询请求,查询请求包括第一文本;响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度;基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本;基于至少一个第二文本,生成针对查询请求的响应。本公开的实施例可以有效的解决文本过长导致语义分散的问题,也能够提高针对大于阈值的文本的检索的准确性。

Description

请求处理方法、装置、设备和存储介质 技术领域
本公开的示例实施例总体涉及计算机领域,特别地涉及请求处理方法、装置、设备和计算机可读存储介质。
背景技术
文本检索是自然语言处理最基本且最重要的任务之一。文本检索目前已经在智能问答、意图识别、语义理解、语义生成等许多场景落地。比如,电子设备可以基于查询请求中包括的文本内容,检索与该文本匹配的其他文本等等。
发明内容
在本公开的第一方面,提供了一种请求处理方法。该方法包括:获取查询请求,查询请求包括第一文本;响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度;基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本;以及基于至少一个第二文本,生成针对查询请求的响应。
在本公开的第二方面,提供了一种用于请求处理的装置。该装置包括:获取模块,被配置为获取查询请求,查询请求包括第一文本;第一生成模块,被配置为响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度;第一确定模块,被配置为基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本;以及第二生成模块,被配置为基于至少一个第二文本,生成针对查询请求的响应。
在本公开的第三方面,提供了一种电子设备。该设备包括至少一 个处理单元;以及至少一个存储器,至少一个存储器被耦合到至少一个处理单元并且存储用于由至少一个处理单元执行的指令。指令在由至少一个处理单元执行时使设备执行第一方面的方法。
在本公开的第四方面,提供了一种计算机可读存储介质。该计算机可读存储介质上存储有计算机程序,计算机程序可由处理器执行以实现第一方面的方法。
应当理解,本内容部分中所描述的内容并非旨在限定本公开的实施例的关键特征或重要特征,也不用于限制本公开的范围。本公开的其它特征将通过以下的描述而变得容易理解。
附图说明
结合附图并参考以下详细说明,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。在附图中,相同或相似的附图标记表示相同或相似的元素,其中:
图1示出了本公开的实施例能够在其中实现的示例环境的示意图;
图2出了根据本公开的一些实施例的请求处理的过程的流程图;
图3示出了根据本公开的一些实施例的请求处理的过程示意图;
图4示出了根据本公开的某些实施例的用于请求处理的装置的示意性结构框图;以及
图5示出了能够实施本公开的多个实施例的电子设备的框图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中示出了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例,相反,提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
需要注意的是,本文中所提供的任何节/子节的标题并不是限制性 的。本文通篇描述了各种实施例,并且任何类型的实施例都可以包括在任何节/子节下。此外,在任一节/子节中描述的实施例可以以任何方式与同一节/子节和/或不同节/子节中描述的任何其他实施例相结合。
在本公开的实施例的描述中,术语“包括”及其类似用语应当理解为开放性包含,即“包括但不限于”。术语“基于”应当理解为“至少部分地基于”。术语“一个实施例”或“该实施例”应当理解为“至少一个实施例”。术语“一些实施例”应当理解为“至少一些实施例”。下文还可能包括其他明确的和隐含的定义。术语“第一”、“第二”等可以指代不同的或相同的对象。下文还可能包括其他明确的和隐含的定义。
本公开的实施例中可能涉及用户的数据、数据的获取和/或使用等。这些方面均遵循相应的法律法规及相关规定。在本公开的实施例中,所有数据的采集、获取、处理、加工、转发、使用等,都是在用户知晓并且确认的前提下进行的。相应地,在实现本公开的各实施例时,均应根据相关法律法规通过适当的方式,将可能所涉及的数据或信息的类型、使用范围、使用场景等告知用户并获得用户的授权。具体的告知和/或授权方式可以根据实际情况和应用场景而变化,本公开的范围在此方面不受限制。
本说明书及实施例中方案,如涉及个人信息处理,则均会在具备合法性基础(例如征得个人信息主体同意,或者为履行合同所必需等)的前提下进行处理,且仅会在规定或者约定的范围内进行处理。用户拒绝处理基本功能所需必要信息以外的个人信息,不会影响用户使用基本功能。
传统的,文本匹配的方法主要包括:基于监督学习的主题提取匹配方法、基于文本分词的匹配的方法,基于文本语义分割后进行匹配的方法、基于语义深度匹配的方法等。这些传统的文本检索的方法仅仅考虑了文本本身的字词/语义信息,没有考虑不恰当的分割会导致语义分割或者语义分散导致的文本检索的准确度不高的问题。
本公开的实施例提出了一种请求处理方案。根据该方案,可以获取查询请求,查询请求包括第一文本;进一步地,可以响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度;进一步地,可以基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本;进一步地,可以基于至少一个第二文本,生成针对查询请求的响应。
基于这样的方式,本公开的实施例可以基于第一文本生成长度小于第一文本的多个关联文本,并基于多个关联文本匹配关联的至少一个第二文本,可以有效的解决因文本过长导致语义分散的问题,也能够提高具有较大长度的文本的检索准确性。
示例环境
图1示出了本公开的实施例能够在其中实现的示例环境100的示意图。如图1所示,示例环境100可以包括电子设备110。
在该示例环境100中,电子设备110可以基于查询请求中的第一文本检索匹配的其他文本,其中第一文本对应的第一长度大于阈值(在本公开中例如也可以称为“长文本”,或者“长文本内容”)。查询请求可以为任意适当的请求,其可以为用户输入的,还可以为电子设备自动生成的,在此不做赘述。
作为一种示例,电子设备可以在接收到用户输入的查询与小说A类似的其他小说的查询请求时,检索与小说A类似的小说B并推荐。
作为另一种示例,电子设备可以在用户基于机器人程序(bot)创建平台创建bot时,基于与bot相关的信息自动为bot推荐工具,其中与bot相关的信息可以为bot的标识、bot的描述、bot的系统提示项等文本信息构成的长文本内容,工具的描述信息相当于基于与bot相关的信息检索到的其他文本。
电子设备110可以是任意类型的移动终端、固定终端或便携式终端,包括移动手机、台式计算机、膝上型计算机、笔记本计算机、上 网本计算机、平板计算机、媒体计算机、多媒体平板、掌上电脑、便携式游戏终端、VR/AR设备、个人通信系统(Personal Communication System,PCS)设备、个人导航设备、个人数字助理(Personal Digital Assistant,PDA)、音频/视频播放器、数码相机/摄像机、定位设备、电视接收器、无线电广播接收器、电子书设备、游戏设备或者前述各项的任意组合,包括这些设备的配件和外设或者其任意组合。在一些实施例中,电子设备110也能够支持任意类型的针对目标用户的接口(诸如“可佩戴”电路等)。
电子设备110也可以是独立的物理服务器,也可以是多个物理服务器构成的服务器集群或者分布式系统,还可以是提供云服务、云数据库、云计算、云函数、云存储、网络服务、云通信、中间件服务、域名服务、安全服务、内容分发网络、以及大数据和人工智能平台等基础云计算服务的云服务器。电子设备110例如可以包括计算系统/服务器,诸如大型机、边缘计算节点、云环境中的计算设备,等等。
应当理解,仅出于示例性的目的描述环境100中各个元素的结构和功能,而不暗示对于本公开的范围的任何限制。
以下将继续参考附图描述本公开的一些示例实施例。
示例过程
图2示出了根据本公开的一些实施例的请求处理的过程200的流程图。过程200可以被实现在电子设备110处。下面参考图1描述过程200。
在框210,电子设备110可以获取查询请求,查询请求包括第一文本。
在一些实施例中,第一文本可以为任意适当的文本,其可以对应于用于查询的基础文本。查询请求可以为任意适当的请求,其可以为用户输入的,还可以为电子设备110自动生成的,在此不做赘述。电子设备110基于查询请求生成针对查询请求的响应的过程中包括基于第一文本进行检索/匹配。
在框220,电子设备110可以响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度。
为了解决文本的长度过长而导致语义分散,进而影响文本检索的准确率的问题,在一些实施例中,电子设备110可以基于第一文本,生成第一模型的第一输入信息,第一输入信息包括第一提示项,第一提示项用于指示第一模型基于第一文本生成预设长度的关联文本。预定长度可以为任意适当的长度,且这个预定长度可以为一个预定的长度数值,还可以为预定的长度范围,比如,第一提示项可以指示第一模型基于第一文本生成长度不大于预定长度阈值的关联文本,还可以指示第一模型基于第一文本生成长度为A的多个关联文本等等。在本公开中,关联文本也可以视为“短文本”或“短文本内容”。
在一些实施例中,第一提示项还可以用于指示第一模型生成的关联文本的数量、关联文本的格式、怎样基于第一文本生成关联文本等等,在此不做赘述。比如,第一提示项可以为“给你一个文本,这个文本中可能包含多个意图,你需要理解这个文本的这些意图,并拆分为可以代表这些意图的较短文本”。
电子设备110可以将第一输入信息输入到第一模型中,以获取由第一模型输出的多个关联文本。这多个关联文本的第二长度小于第一文本的第一长度,且这多个关联文本对应的第二长度可以相同,也可以不同,具体可以根据需求设置。
在一些实施例中,第一模型可以为任意适当的模型。为了提高检索的准确性,第一模型可以为具有强化泛化能力,且能较好的处理长距离依赖的模型,比如可以为语言模型等。
在一些实施例中,电子设备110可以利用第一模型对第一文本进行语义理解,并基于语义理解结果生成多个关联文本。
在一些实施例中,电子设备110还可以利用第一模型的世界知识,对第一文本进行语义理解,来生成多个关联文本,其中世界知识可以为任意适当且通用的可协助第一模型理解文本语义的知识。比如,世 界知识可以包括“世界上有七大洲”等等。
电子设备110利用第一模型来生成多个关联文本可以有效的解决因文本长度过程而导致语义分散发生灾难性遗忘的问题,且可以提高生成的多个关联文本与第一文本的关联性,并能够进行精确语义分割,做到语义上的强泛化性。
在框230,电子设备110可以基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本。
在一些实施例中,第二文本的长度可以大于多个关联文本的长度,且第二文本的长度可以大于预定的阈值,即这个第二文本可以为与第一文本匹配的文本(例如,长文本内容)。
在一些实施例中,电子设备110可以基于多个关联文本的语义信息从一组候选文本中确定与语义信息匹配的至少一个第二文本,即电子设备110可以基于多个关联文本的语义信息以及这一组候选文本的语义信息的比较结果,从一组候选文本中确定与语义信息匹配的至少一个第二文本。
在另一些实施例中,电子设备110可以基于多个关联文本的文本信息从一组候选文本中确定与文本信息匹配的至少一个第二文本,即电子设备110可以基于多个关联文本的文本信息以及这一组候选文本的文本信息的比较结果,从一组候选文本中确定与文本信息匹配的至少一个第二文本。文本信息可以为多个关联文本本身,还可以为基于关联文本提取的其他的文本信息,比如关键字等等。
在另一些实施例中,电子设备110可以基于多个关联文本的文本信息和语义信息从一组候选文本中确定与文本信息或/和语义信息匹配的至少一个第二文本。在一些实施例中,电子设备110可以从一组候选文本中确定与文本信息匹配的至少一个第一候选文本。电子设备110可以从一组候选文本中确定与语义信息匹配的至少一个第二候选文本。电子设备110可以基于至少一个第一候选文本以及至少一个第二候选文本,确定与查询请求匹配的至少一个第二文本。作为一种示例,电子设备110可以将这至少一个第一候选文本以及这至少一个第 二候选文本均确定为第二文本。作为另一种示例,电子设备110可以将这至少一个第一候选文本以及这至少一个第二候选文本中的公共候选文本确定为第二文本。比如,一组候选文本包括候选文本A、候选文本B、候选文本C以及候选文本D,电子设备110基于语义信息从一组候选文本中确定与查询请求匹配的候选文本A与候选文本B,电子设备110基于文本信息从一组候选文本中确定候选文本A与候选文本C,那么电子设备110可以确定候选文本A为第二文本。
在一些实施例中,电子设备110可以基于至少一个第一候选文本以及至少一个第二候选文本与多个关联文本的关联性,确定至少一个第一候选文本以及至少一个第二候选文本对应的第一排序结果。电子设备110可以基于经排序的多个候选文本,确定与查询请求匹配的至少一个第二文本。
在一些实施例中,电子设备可以将多个关联文本对应的特征表示进行拼接,获取目标特征表示,并基于至少一个第一候选文本以及至少一个第二候选文本与多个关联文本对应的特征表示以及目标特征表示,确定至少一个第一候选文本以及至少一个第二候选文本与多个关联文本的关联性的相关性。
在另一些实施例中,电子设备110可以基于至少一个第一候选文本以及多个第二候选文本与第一文本的关联性,确定至少一个第一候选文本以及多个第二候选文本对应的第二排序结果。电子设备110可以基于经排序的多个候选文本,确定与查询请求匹配的至少一个第二文本。
作为一种示例,电子设备110可以基于各个候选文本对应的语义特征表示与第一文本对应的语义特征表示,确定目标欧式距离。进一步地,电子设备110可以基于目标欧式距离确定各个候选文本与第一文本的相关性,其中目标欧式距离越小,相关性越高。
在一些实施例中,电子设备110可以基于至少一个第一候选文本以及多个第二候选文本的属性信息,确定与查询请求匹配的至少一个第二文本。在一些实施例中,属性信息可以为任意适当的信息,比如 知名度、用户反馈信息等等。以第二文本为小说为例,用户反馈信息可以为用户查看的次数、用户点赞的次数、用户收藏的次数等等。
下面针对这一组候选文本对应的语义信息和文本信息的确定过程进行说明。
在一些实施例中,电子设备110可以获取一组候选文本。电子设备110可以通过将一组候选文本输入到第二模型中,获取第二模型输出的一组候选文本中每个候选文本对应的多个描述文本,其中描述文本对应的长度小于候选文本对应的长度。在一些实施例中,第二模型可以为任意适当的模型。第二模型可以为具有强化泛化能力,且能较好的处理长文本内容的模型。比如可以为语言模型等等。
在一些实施例中,电子设备110可以利用第二模型对每个候选文本进行语义理解,并基于语义理解结果生成多个描述文本。在一些实施例中,电子设备110还可以利用第二模型的世界知识,对每个候选文本进行语义理解,来生成多个描述文本。
电子设备110可以确定一组候选文本中每个候选文本对应的多个描述文本的语义信息。电子设备110还可以确定一组候选文本中每个候选文本对应的多个描述文本的文本信息。
电子设备110还可以在向量数据库中保存候选文本、候选文本对应的描述文本、候选文本对应的描述文本的语义信息以及候选文本对应的描述文本的文本信息的对应关系,以支持电子设备110可以基于查询请求中的第一文本生成的多个关联文本的语义信息和/或文本信息,匹配或检索到的与查询请求匹配的至少一个第二文本。
在一些实施例中,针对基于长文本所生成的每个关联文本,电子设备110可以确定多个与该关联文本在语义信息和/或文本信息上匹配的一项或多项描述文本。进一步地,电子设备110可以确定与该一项或多项描述文本对应的候选文本,以作为与长文本匹配的第二文本。
在框240,电子设备110可以基于至少一个第二文本,生成针对查询请求的响应。
在一些实施例中,电子设备110可以直接基于这至少一个第二文 本,生成针对查询请求的响应,即将这至少一个第二文本均确定为与第一文本匹配的文本。
在另外一些实施例中,电子设备110还可以基于这至少一个第二文本中的任意一个作为目标文本,并基于目标文本生成针对查询请求的响应。
在另一些实施例中,电子设备110还可以向第三模型提供第一文本、多个关联文本以及至少一个第二文本。电子设备110可以获取由第三模型从至少一个第二文本中确定的目标文本,目标文本基于第一文本和多个关联文本所确定。目标文本为电子设备110基于第一文本以及多个关联文本进一步从至少一个第二文本中确定的与第一文本更相关的文本。电子设备110可以利用目标文本,生成针对查询请求的响应。
在一些实施例中,第三模型可以被指示为:基于至少一个第二文本与第一文本和多个关联文本的相关性,从至少一个第二文本中确定目标文本。作为一种示例,电子设备110可以基于第一文本、多个关联文本以及至少一个第二文本,生成输入到第三模型中的第二输入信息,其中,第二输入信息包括第二提示项。
第二提示项可以用于指示第三模型怎样从至少一个第二文本中确定目标文本等等,在此不做赘述。比如,第二提示项可以为“给你一个长文本以及与这个长文本对应的一些意图,这些意图并不一定覆盖文本对应的所有意图。另外给你一些候选文本,你需要理解这个文本,并根据这个文本以及这个文本对应的多个意图从这些候选文本中找出与这个长文本匹配的文本”。
电子设备110可以利用这个第二输入信息,获取第三模型基于这个第二输入信息确定的目标文本。
图3示出了根据本公开的一些实施例的请求处理的过程示意图,现针对图3进行说明。
电子设备110可以获取查询请求,查询请求中包括长度大于阈值的长文本。电子设备110可以利用语言模型理解长文本对应的语义信 息,生成m个长度小于等于小阈值的短文本。电子设备110可以获取这m个短文本对应的语义表征(语义信息)。电子设备110可以基于m个长文本以及向量数据库中保存的各个候选文本分别对应的多个描述文本,确定m*k个第一匹配文本。电子设备110还可以基于m个长文本对应的语义表征以及向量数据库中保存的各个候选文本分别对应的多个描述文本的语义表征,确定m*k个第一匹配文本。
需要说明的是,向量数据库中的各个候选文本分别对应的多个描述文本也可以为基于语言模型理解各个候选文本的语义信息后获取的,且各个候选文本分别对应的多个描述文本的语义表征的获取方式也可以与长文本对应的m个短文本对应的语义表征的获取方式相同,也可以不同,在此不做赘述。
在基于短文本的文本匹配和短文本的语义匹配确定出(m+m)*k个第一匹配文本后,电子设备110可以从(m+m)*k个匹配文本中,基于模型筛选出a个第二匹配文本,其中a小于(m+m)*k。进一步地,电子设备110可以利用语言模型,基于查询请求中的长文本、m个短文本以及a个第二匹配文本进一步进行匹配,获取与长文本匹配的目标匹配文本,以基于目标匹配文本生成针对查询请求的响应。
需要说明的是,电子设备110可以预先配置参数配置信息,参数配置信息可以包括但是不限于上述的m、k以及a等等。
基于这样的方式,本公开的实施例可以基于第一文本生成长度小于第一文本的多个关联文本,并基于多个关联文本检索与第一文本关联的至少一个第二文本,可以有效的解决因文本过长导致语义分散的问题,也能够提高具有较大长度的文本的检索的准确性。
示例装置和设备
本公开的实施例还提供了用于实现上述方法或过程的相应装置。图4示出了根据本公开的某些实施例的用于请求处理的装置400的示意性结构框图。装置400可以被实现为或者被包括在如上文所讨论的电子设备110中。装置400中的各个模块/组件可以由硬件、软件、固 件或者它们的任意组合来实现。
如图4所示,装置400包括获取模块410,被配置为获取查询请求,查询请求包括第一文本;第一生成模块420,被配置为响应于第一文本的第一长度大于阈值,基于第一文本生成多个关联文本,关联文本的第二长度小于第一长度;第一确定模块430,被配置为基于多个关联文本的多个第一特征表示,从一组候选文本中确定与查询请求匹配的至少一个第二文本;以及第二生成模块440,被配置为基于至少一个第二文本,生成针对查询请求的响应。
在一些实施例中,第一生成模块420,具体被配置为基于第一文本,生成第一模型的第一输入信息,第一输入信息包括第一提示项,第一提示项用于指示第一模型基于第一文本生成预设长度的关联文本;以及获取由第一模型输出的多个关联文本。
在一些实施例中,第一确定模块420,具体被配置为基于多个关联文本的语义信息和/或文本信息,从一组候选文本中确定与文本信息和/或语义信息匹配的至少一个第二文本。
在一些实施例中,第一确定模块420,具体被配置为从一组候选文本中确定与文本信息匹配的至少一个第一候选文本;从一组候选文本中确定与语义信息匹配的至少一个第二候选文本;以及基于至少一个第一候选文本以及至少一个第二候选文本,确定与查询请求匹配的至少一个第二文本。
在一些实施例中,装置400还包括语义确定模块,被配置为:获取一组候选文本;通过将一组候选文本输入到第二模型中,获取第二模型输出的一组候选文本中每个候选文本对应的多个描述文本,其中描述文本的长度小于候选文本;以及基于多个描述文本,确定相应候选文本的语义信息。
在一些实施例中,第一确定模块420,具体被配置为基于各候选文本与多个关联文本和/或第一文本的关联性,排序至少一个第一候选文本和至少一个第二候选文本;以及基于经排序的候选文本,确定与查询请求匹配的至少一个第二文本。
在一些实施例中,第一确定模块420,具体被配置为基于各候选文本的属性信息,从至少一个第一候选文本和至少一个第二候选文本中确定与查询请求匹配的至少一个第二文本。
在一些实施例中,第二生成模块440,具体被配置为向第三模型提供第一文本、多个关联文本以及至少一个第二文本;
获取由第三模型从至少一个第二文本中确定的目标文本,目标文本基于第一文本和多个关联文本所确定;以及利用目标文本,生成针对查询请求的响应。
在一些实施例中,第三模型被指示为:基于至少一个第二文本与第一文本和多个关联文本的相关性,从至少一个第二文本中确定目标文本。
装置400中所包括的单元可以利用各种方式来实现,包括软件、硬件、固件或其任意组合。在一些实施例中,一个或多个单元可以使用软件和/或固件来实现,例如存储在存储介质上的机器可执行指令。除了机器可执行指令之外或者作为替代,装置400中的部分或者全部单元可以至少部分地由一个或多个硬件逻辑组件来实现。作为示例而非限制,可以使用的示范类型的硬件逻辑组件包括现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准品(ASSP)、片上系统(SOC)、复杂可编程逻辑器件(CPLD),等等。
图5示出了其中可以实施本公开的一个或多个实施例的电子设备500的框图。应当理解,图5所示出的电子设备500仅仅是示例性的,而不应当构成对本文所描述的实施例的功能和范围的任何限制。图5所示出的电子设备500可以用于实现图1所示的电子设备110。
如图5所示,电子设备500是通用电子设备的形式。电子设备500的组件可以包括但不限于一个或多个处理器或处理单元510、存储器520、存储设备530、一个或多个通信单元540、一个或多个输入设备550以及一个或多个输出设备560。处理单元510可以是实际或虚拟处理器并且能够根据存储器520中存储的程序来执行各种处理。在多处理器系统中,多个处理单元并行执行计算机可执行指令,以提高电 子设备500的并行处理能力。
电子设备500通常包括多个计算机存储介质。这样的介质可以是电子设备500可访问的任何可以获取的介质,包括但不限于易失性和非易失性介质、可拆卸和不可拆卸介质。存储器520可以是易失性存储器(例如寄存器、高速缓存、随机访问存储器(RAM))、非易失性存储器(例如,只读存储器(ROM)、电可擦除可编程只读存储器(EEPROM)、闪存)或它们的某种组合。存储设备530可以是可拆卸或不可拆卸的介质,并且可以包括机器可读介质,诸如闪存驱动、磁盘或者任何其他介质,其可以能够用于存储信息和/或数据(例如用于训练的训练数据)并且可以在电子设备500内被访问。
电子设备500可以进一步包括另外的可拆卸/不可拆卸、易失性/非易失性存储介质。尽管未在图5中示出,可以提供用于从可拆卸、非易失性磁盘(例如“软盘”)进行读取或写入的磁盘驱动和用于从可拆卸、非易失性光盘进行读取或写入的光盘驱动。在这些情况中,每个驱动可以由一个或多个数据介质接口被连接至总线(未示出)。存储器520可以包括计算机程序产品525,其具有一个或多个程序模块,这些程序模块被配置为执行本公开的各种实施例的各种方法或动作。
通信单元540实现通过通信介质与其他电子设备进行通信。附加地,电子设备500的组件的功能可以以单个计算集群或多个计算机器来实现,这些计算机器能够通过通信连接进行通信。因此,电子设备500可以使用与一个或多个其他服务器、网络个人计算机(PC)或者另一个网络节点的逻辑连接来在联网环境中进行操作。
输入设备550可以是一个或多个输入设备,例如鼠标、键盘、追踪球等。输出设备560可以是一个或多个输出设备,例如显示器、扬声器、打印机等。电子设备500还可以根据需要通过通信单元540与一个或多个外部设备(未示出)进行通信,外部设备诸如存储设备、显示设备等,与一个或多个使得用户与电子设备500交互的设备进行通信,或者与使得电子设备500与一个或多个其他电子设备通信的任 何设备(例如,网卡、调制解调器等)进行通信。这样的通信可以经由输入/输出(I/O)接口(未示出)来执行。
根据本公开的示例性实现方式,提供了一种计算机可读存储介质,其上存储有计算机可执行指令,其中计算机可执行指令被处理器执行以实现上文描述的方法。根据本公开的示例性实现方式,还提供了一种计算机程序产品,计算机程序产品被有形地存储在非瞬态计算机可读介质上并且包括计算机可执行指令,而计算机可执行指令被处理器执行以实现上文描述的方法。
这里参照根据本公开实现的方法、装置、设备和计算机程序产品的流程图和/或框图描述了本公开的各个方面。应当理解,流程图和/或框图的每个方框以及流程图和/或框图中各方框的组合,都可以由计算机可读程序指令实现。
这些计算机可读程序指令可以提供给通用计算机、专用计算机或其他可编程数据处理装置的处理单元,从而生产出一种机器,使得这些指令在通过计算机或其他可编程数据处理装置的处理单元执行时,产生了实现流程图和/或框图中的一个或多个方框中规定的功能/动作的装置。也可以把这些计算机可读程序指令存储在计算机可读存储介质中,这些指令使得计算机、可编程数据处理装置和/或其他设备以特定方式工作,从而,存储有指令的计算机可读介质则包括一个制造品,其包括实现流程图和/或框图中的一个或多个方框中规定的功能/动作的各个方面的指令。
可以把计算机可读程序指令加载到计算机、其他可编程数据处理装置、或其他设备上,使得在计算机、其他可编程数据处理装置或其他设备上执行一系列操作步骤,以产生计算机实现的过程,从而使得在计算机、其他可编程数据处理装置、或其他设备上执行的指令实现流程图和/或框图中的一个或多个方框中规定的功能/动作。
附图中的流程图和框图显示了根据本公开的多个实现的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段或指令的一部 分,模块、程序段或指令的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个连续的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或动作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
以上已经描述了本公开的各实现,上述说明是示例性的,并非穷尽性的,并且也不限于所公开的各实现。在不偏离所说明的各实现的范围和精神的情况下,对于本技术领域的普通技术人员来说许多修改和变更都是显而易见的。本文中所用术语的选择,旨在最好地解释各实现的原理、实际应用或对市场中的技术的改进,或者使本技术领域的其他普通技术人员能理解本文公开的各个实现方式。

Claims (12)

  1. 一种请求处理方法,包括:
    获取查询请求,所述查询请求包括第一文本;
    响应于所述第一文本的第一长度大于阈值,基于所述第一文本生成多个关联文本,关联文本的第二长度小于所述第一长度;
    基于所述多个关联文本的多个第一特征表示,从一组候选文本中确定与所述查询请求匹配的至少一个第二文本;以及
    基于所述至少一个第二文本,生成针对所述查询请求的响应。
  2. 根据权利要求1所述的方法,其中基于所述第一文本生成多个关联文本包括:
    基于所述第一文本,生成第一模型的第一输入信息,所述第一输入信息包括第一提示项,所述第一提示项用于指示所述第一模型基于所述第一文本生成预设长度的关联文本;以及
    获取由所述第一模型输出的所述多个关联文本。
  3. 根据权利要求1所述的方法,其中基于所述多个关联文本的多个第一特征表示,从一组候选文本中确定与所述查询请求匹配的至少一个第二文本包括:
    基于所述多个关联文本的语义信息和/或文本信息,从一组候选文本中确定与所述文本信息和/或所述语义信息匹配的所述至少一个第二文本。
  4. 根据权利要求3所述的方法,其中基于所述多个关联文本的语义信息和/或文本信息,从一组候选文本中确定与所述文本信息和/或所述语义信息匹配的所述至少一个第二文本包括:
    从所述一组候选文本中确定与所述文本信息匹配的至少一个第一候选文本;
    从所述一组候选文本中确定与所述语义信息匹配的至少一个第二候选文本;以及
    基于所述至少一个第一候选文本以及所述至少一个第二候选文 本,确定与所述查询请求匹配的所述至少一个第二文本。
  5. 根据权利要求4所述的方法,还包括:
    获取所述一组候选文本;
    通过将所述一组候选文本输入到第二模型中,获取所述第二模型输出的所述一组候选文本中每个候选文本对应的多个描述文本,其中描述文本的长度小于候选文本;以及
    基于所述多个描述文本,确定相应候选文本的语义信息。
  6. 根据权利要求4所述的方法,其中基于所述至少一个第一候选文本以及所述至少一个第二候选文本,确定与所述查询请求匹配的所述至少一个第二文本包括:
    基于各候选文本与所述多个关联文本和/或所述第一文本的关联性,排序所述至少一个第一候选文本和所述至少一个第二候选文本;以及
    基于经排序的候选文本,确定与所述查询请求匹配的所述至少一个第二文本。
  7. 根据权利要求4所述的方法,其中基于所述至少一个第一候选文本以及所述至少一个第二候选文本,确定与所述查询请求匹配的所述至少一个第二文本包括:
    基于各候选文本的属性信息,从所述至少一个第一候选文本和所述至少一个第二候选文本中确定与所述查询请求匹配的所述至少一个第二文本。
  8. 根据权利要求1所述的方法,其中基于所述至少一个第二文本,生成针对所述查询请求的响应包括:
    向第三模型提供所述第一文本、所述多个关联文本以及所述至少一个第二文本;
    获取由所述第三模型从所述至少一个第二文本中确定的目标文本,所述目标文本基于所述第一文本和所述多个关联文本所确定;以及
    利用所述目标文本,生成针对所述查询请求的响应。
  9. 根据权利要求8所述的方法,其中所述第三模型被指示为:基于所述至少一个第二文本与所述第一文本和所述多个关联文本的相关性,从所述至少一个第二文本中确定所述目标文本。
  10. 一种用于请求处理的装置,包括:
    获取模块,被配置为获取查询请求,所述查询请求包括第一文本;
    第一生成模块,被配置为响应于所述第一文本的第一长度大于阈值,基于所述第一文本生成多个关联文本,关联文本的第二长度小于所述第一长度;
    第一确定模块,被配置为基于所述多个关联文本的多个第一特征表示,从一组候选文本中确定与所述查询请求匹配的至少一个第二文本;以及
    第二生成模块,被配置为基于所述至少一个第二文本,生成针对所述查询请求的响应。
  11. 一种电子设备,包括:
    至少一个处理单元;以及
    至少一个存储器,所述至少一个存储器被耦合到所述至少一个处理单元并且存储用于由所述至少一个处理单元执行的指令,所述指令在由所述至少一个处理单元执行时使所述电子设备执行根据权利要求1至9中任一项所述的方法。
  12. 一种计算机可读存储介质,其上存储有计算机程序,所述计算机程序被处理器执行时实现根据权利要求1至9中任一项所述的方法。
PCT/CN2024/099090 2024-06-13 2024-06-13 请求处理方法、装置、设备和存储介质 Pending WO2025255789A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
CN202480004209.7A CN121532760A (zh) 2024-06-13 2024-06-13 请求处理方法、装置、设备和存储介质
PCT/CN2024/099090 WO2025255789A1 (zh) 2024-06-13 2024-06-13 请求处理方法、装置、设备和存储介质

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/CN2024/099090 WO2025255789A1 (zh) 2024-06-13 2024-06-13 请求处理方法、装置、设备和存储介质

Publications (1)

Publication Number Publication Date
WO2025255789A1 true WO2025255789A1 (zh) 2025-12-18

Family

ID=98049939

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2024/099090 Pending WO2025255789A1 (zh) 2024-06-13 2024-06-13 请求处理方法、装置、设备和存储介质

Country Status (2)

Country Link
CN (1) CN121532760A (zh)
WO (1) WO2025255789A1 (zh)

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112417885A (zh) * 2020-11-17 2021-02-26 平安科技(深圳)有限公司 基于人工智能的答案生成方法、装置、计算机设备及介质
CN114385777A (zh) * 2022-01-14 2022-04-22 平安科技(深圳)有限公司 文本数据处理方法、装置、计算机设备和存储介质
US20220300543A1 (en) * 2021-06-15 2022-09-22 Beijing Baidu Netcom Science Technology Co., Ltd. Method of retrieving query, electronic device and medium
CN115835178A (zh) * 2022-10-11 2023-03-21 光宝科技股份有限公司 在服务管理与编排中通讯的方法及系统
CN116108140A (zh) * 2023-02-14 2023-05-12 南阳理工学院 一种自然语言匹配法律条文的方法

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN112417885A (zh) * 2020-11-17 2021-02-26 平安科技(深圳)有限公司 基于人工智能的答案生成方法、装置、计算机设备及介质
US20220300543A1 (en) * 2021-06-15 2022-09-22 Beijing Baidu Netcom Science Technology Co., Ltd. Method of retrieving query, electronic device and medium
CN114385777A (zh) * 2022-01-14 2022-04-22 平安科技(深圳)有限公司 文本数据处理方法、装置、计算机设备和存储介质
CN115835178A (zh) * 2022-10-11 2023-03-21 光宝科技股份有限公司 在服务管理与编排中通讯的方法及系统
CN116108140A (zh) * 2023-02-14 2023-05-12 南阳理工学院 一种自然语言匹配法律条文的方法

Also Published As

Publication number Publication date
CN121532760A (zh) 2026-02-13

Similar Documents

Publication Publication Date Title
CN112115232B (zh) 一种数据纠错方法、装置及服务器
CN111797210A (zh) 基于用户画像的信息推荐方法、装置、设备及存储介质
US11120214B2 (en) Corpus generating method and apparatus, and human-machine interaction processing method and apparatus
CN106462807A (zh) 根据大规模非结构化数据学习多媒体语义
WO2025189617A1 (zh) 请求处理的方法、装置、设备和存储介质
CN110569289A (zh) 基于大数据的列数据处理方法、设备及介质
CN117251879A (zh) 基于信任扩展的安全存储与查询方法、系统及计算机储存介质
CN103891244B (zh) 一种进行数据存储和检索的方法及装置
WO2025246492A1 (zh) 信息处理方法、装置、设备和存储介质
WO2025223030A1 (zh) 信息处理方法、装置、设备和存储介质
WO2025223031A1 (zh) 信息处理方法、装置、设备和存储介质
WO2025255789A1 (zh) 请求处理方法、装置、设备和存储介质
WO2025222808A1 (zh) 代码编辑的方法、装置、设备和存储介质
CN119557421A (zh) 基于关键字的搜索方法、装置、介质及设备
WO2025077303A1 (zh) 信息处理的方法、装置、设备和存储介质
CN111695031A (zh) 基于标签的搜索方法、装置、服务器及存储介质
CN110674383A (zh) 舆情查询方法、装置及设备
CN115934900A (zh) 单词的派生联想方法、装置、计算机设备和存储介质
WO2025175778A1 (zh) 代码检索的方法、装置、设备和存储介质
US20260057025A1 (en) Content query
CN120910049B (zh) 一种数据的补全方法、装置、电子设备及介质
WO2025081893A1 (zh) 信息搜索的方法、装置、设备和存储介质
CN113934817B (zh) 目标店铺的查询方法、装置、计算机设备和存储介质
CN112084290A (zh) 一种数据检索方法、装置、设备及存储介质
WO2025119362A1 (zh) 信息搜索方法、装置、设备和存储介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 24942961

Country of ref document: EP

Kind code of ref document: A1