WO2022174669A1 - 信息生成方法、装置、电子设备和计算机可读介质 - Google Patents

信息生成方法、装置、电子设备和计算机可读介质 Download PDF

Info

Publication number
WO2022174669A1
WO2022174669A1 PCT/CN2021/140219 CN2021140219W WO2022174669A1 WO 2022174669 A1 WO2022174669 A1 WO 2022174669A1 CN 2021140219 W CN2021140219 W CN 2021140219W WO 2022174669 A1 WO2022174669 A1 WO 2022174669A1
Authority
WO
WIPO (PCT)
Prior art keywords
graph model
information
parameter
target item
mentioned
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2021/140219
Other languages
English (en)
French (fr)
Inventor
梁嘉辉
鲍军威
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Beijing Jingdong Century Trading Co Ltd
Beijing Wodong Tianjun Information Technology Co Ltd
Original Assignee
Beijing Jingdong Century Trading Co Ltd
Beijing Wodong Tianjun Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Beijing Jingdong Century Trading Co Ltd, Beijing Wodong Tianjun Information Technology Co Ltd filed Critical Beijing Jingdong Century Trading Co Ltd
Priority to JP2023550238A priority Critical patent/JP7656057B2/ja
Priority to US18/277,928 priority patent/US12306858B2/en
Publication of WO2022174669A1 publication Critical patent/WO2022174669A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/205Parsing
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/332Query formulation
    • G06F16/3325Reformulation based on results of preceding query
    • G06F16/3326Reformulation based on results of preceding query using relevance feedback from the user, e.g. relevance feedback on documents, documents sets, document terms or passages
    • G06F16/3328Reformulation based on results of preceding query using relevance feedback from the user, e.g. relevance feedback on documents, documents sets, document terms or passages using graphical result space presentation or visualisation
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/901Indexing; Data structures therefor; Storage structures
    • G06F16/9024Graphs; Linked lists
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/33Querying
    • G06F16/3331Query processing
    • G06F16/334Query execution
    • G06F16/3347Query execution using vector based model
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/34Browsing; Visualisation therefor
    • G06F16/345Summarisation for human users
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/36Creation of semantic tools, e.g. ontology or thesauri
    • G06F16/367Ontology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/38Retrieval characterised by using metadata, e.g. metadata not derived from the content or metadata generated manually
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/42Data-driven translation
    • G06F40/44Statistical methods, e.g. probability models
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/40Processing or translation of natural language
    • G06F40/53Processing of non-Latin text
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods

Definitions

  • the electronic device 101 may first acquire a pre-built graphical model 105 associated with the target item category, serialized text data 104 and structured text data 103 related to the target item 102 .
  • the above-mentioned graph model 105 represents the relationship between the parameter name of the target item category and each parameter content information
  • the nodes in the above-mentioned graph model 105 represent the parameter name or parameter content information of the above-mentioned target item category, in the above-mentioned graph model 105
  • the value corresponding to the edge of represents the degree of association information between the parameter name and the parameter content information.
  • the above graph model 105 includes: "screen size -> 6.5 inches, association degree information: 0.6; screen size -> 5.5 inches, association degree information: 0.4; color -> light red, association degree information: 0.4 ;color->white, correlation degree information: 0.4; color->black, correlation degree information: 0.3, battery capacity->500, correlation degree information: 0.3; battery capacity->mHh, correlation degree information: 0.9; battery capacity- >1000, correlation degree information: 0.6; pixel->500, correlation degree information: 0.5; pixel->1000, correlation degree information: 0.4; pixel->800, correlation degree information: 0.5; pixel->10,000, correlation degree information : 0.8".
  • the above-mentioned target item 102 may be: "mobile phone".
  • the fused graph model 109 may include: "color->light red, correlation degree information: 0.4; battery capacity->1000, correlation degree information: 0.6; pixel->1000, correlation degree information: 0.4 ; screen size -> 5.5 inches, relevance information: 0.4; ** mobile phone; quality; large; good; wear-resistant; high”.
  • the above-mentioned summary information 110 may be: "** mobile phone, good quality, large screen size, 5.5 inches, large battery capacity, 1000mAh, high pixel, 10 million”.
  • the above electronic device 101 may be hardware or software.
  • the electronic device When the electronic device is hardware, it can be implemented as a distributed cluster composed of multiple servers or terminal devices, or can be implemented as a single server or a single terminal device.
  • the electronic device When the electronic device is embodied as software, it can be installed in the hardware devices listed above. It can be implemented, for example, as multiple software or software modules for providing distributed services, or as a single software or software module. There is no specific limitation here.
  • FIG. 1 the number of electronic devices in FIG. 1 is merely illustrative. There may be any number of electronic devices depending on implementation needs.
  • the information generation method includes the following steps:
  • the execution body of the information generation method may acquire a pre-built graph model associated with the target item category, the target item related graph through a wired connection or a wireless connection.
  • serialized text data and structured text data may be acquired.
  • the above-mentioned graph model represents the relationship between the parameter name of the target item category and each parameter content information
  • the nodes in the above-mentioned graph model represent the parameter name or parameter content information of the above-mentioned target item category
  • the edge in the above-mentioned graph model corresponds to
  • the value of represents the degree of association information between the above-mentioned parameter name and the above-mentioned parameter content information.
  • the above-mentioned association degree information may be a value between 0 and 1.
  • wireless connection methods may include but are not limited to 3G/4G/5G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other wireless connection currently known or developed in the future connection method.
  • the above-mentioned execution body may perform word segmentation on the above-mentioned serialized text data to obtain a word set.
  • the above serialized text data may be segmented by stuttering word segmentation to obtain a word set.
  • Step 403 Match the target word in the vocabulary set with the parameter name in the graph model, and generate a matched graph model according to the matching result.
  • the execution body may match the target word in the vocabulary set with the parameter names in the graph model, and generate a matched graph model according to the matching result.
  • the above-mentioned target word may be a word that is associated with the parameter name in the graph model.
  • the above-mentioned execution body can convert the matching result into the form of a graph, and can obtain a matched graph model.
  • a matched graph model is generated according to the above matching result and the target edge in the above graph model.
  • the above-mentioned target edge is an edge in the above-mentioned graph model representing the degree of association information between the above-mentioned target word and the parameter name corresponding to the above-mentioned target word.
  • the above-mentioned execution body may determine the node information of the above-mentioned matched graph model according to the above-mentioned matching result.
  • the target edge in the above-mentioned graph model is used as the edge of the above-mentioned matched graph model, so as to generate the matched graph model.
  • the above-mentioned execution body may generate a merged graphical model according to the above-mentioned structured text data and the above-mentioned matched graphical model.
  • the above-mentioned execution body may integrate the structured text data into the above-mentioned matched graph model to obtain the fused graph model.
  • the fused graph model embodies the structured textual information related to the target item.
  • the above-mentioned execution body may generate the abstract information of the above-mentioned target item according to the above-mentioned fused graph model.
  • the above-mentioned execution body may first convert the above-mentioned fused graph model into a corresponding vector. Then, the above-mentioned corresponding vectors and the vectors corresponding to each node in the above-mentioned fused graph model are input into the pre-trained recurrent neural network to obtain the summary information of the above-mentioned target item.
  • the first step is to acquire item information of each item under the target item category, wherein the item information includes a parameter name and at least one parameter value.
  • each item information in the item information of each item corresponds to at least one parameter name.
  • the parameter name set corresponding to the item information of each item can be determined. There may be multiple duplicate parameter names in the above parameter name set. In this way, the parameter name set is deduplicated, and the deduplicated parameter name set can be obtained.
  • each parameter name in the parameter name set after deduplication corresponds to at least one parameter content information.
  • the information generation methods of some embodiments of the present disclosure can accurately and effectively generate the abstract information of the target item by considering the characteristic information of the serialized text data and the structured text data .
  • recurrent neural networks can only model sequence text information, since item information includes not only unstructured text sequences, but also structured information.
  • item information includes not only unstructured text sequences, but also structured information.
  • the current processing method is still as a sequence of strings. This processing method may lose the structured information, thereby reducing the expressive ability of the model.
  • the information generation methods of some embodiments of the present disclosure may first acquire a pre-built graphical model associated with the target item category, serialized text data and structured text data related to the target item.
  • the above-mentioned graph model represents the relationship between the parameter name of the target item category and each parameter content information
  • the nodes in the above-mentioned graph model represent the parameter name or parameter content information of the above-mentioned target item category
  • the edge in the above-mentioned graph model corresponds to The value of represents the degree of association information between the above-mentioned parameter name and the above-mentioned parameter content information.
  • the above graph model is used to determine the parameter name corresponding to the target word in the subsequent vocabulary set, which lays a foundation for the subsequent generation of the matched graph model. It should be noted that obtaining the above graph model is used to convert the serialized text data related to the target item into structured data. Then, perform word segmentation on the above serialized text data to obtain a word set. Further, the target word in the vocabulary set is matched with the parameter name in the graph model, and a matched graph model is generated according to the matching result. Next, according to the above-mentioned structured text data and the above-mentioned matched graph model, a fused graph model is generated. Here, the fused graph model embodies the structured textual information related to the target item.
  • the summary information of the above-mentioned target item is generated. Therefore, the above information generation method can accurately and effectively generate the summary information of the target item by considering the characteristic information of the serialized text data and the structured text data.
  • Step 501 Acquire a pre-built graph model associated with the target item category, serialized text data and structured text data related to the target item.
  • Step 502 Perform word segmentation on the above serialized text data to obtain a word set.
  • Step 503 Match the target word in the vocabulary set with the parameter names in the graph model, and generate a matched graph model according to the matching result.
  • Step 504 Generate a fused graph model according to the above structured text data and the above matched graph model.
  • the execution body may input the vector matrix and the vector corresponding to each node in the fused graph model into a pre-trained Graph Convolutional Neural Network (GCN, Graph Convolutional Network) to obtain the first output result .
  • GCN Graph Convolutional Neural Network
  • the above graph convolutional network may have multiple layers of graph convolutional layers.
  • the above graph convolution layer can aggregate the neighbor information of each node through the adjacency matrix, thereby enhancing the representation of the current node and improving the accuracy of the model's input modeling.
  • Multi-layer graph convolutional layer modeling can obtain information about multi-hop neighbors.
  • Step 508 Encode the first output result to obtain the second output result.
  • the execution subject may decode the second output result to obtain the summary information of the target item.
  • the above-mentioned execution body may input the above-mentioned second output result to a pre-trained decoding network to obtain the second output result.
  • the process 500 of the information generation method in some embodiments corresponding to FIG. 5 embodies the generation of the above-mentioned target item according to the above-mentioned fused graph model summary information steps. Therefore, the solutions described in these embodiments can generate the summary information of the above-mentioned target item more accurately and effectively.
  • the present disclosure provides some embodiments of an information generating apparatus. These apparatus embodiments correspond to the above-mentioned method embodiments in FIG. 4 , and the apparatus can be specifically applied in various electronic devices.
  • the word segmentation unit 602 is configured to perform word segmentation on the above serialized text data to obtain a word set.
  • the matching generating unit 603 is configured to match the target word in the above-mentioned vocabulary set with the parameter names in the above-mentioned graph model, and generate a matched graph model according to the matching result.
  • the first generating unit 604 is configured to generate a fused graph model according to the above-mentioned structured text data and the above-mentioned matched graph model.
  • the second generating unit 605 is configured to generate the summary information of the target item according to the fused graph model.
  • the second generating unit 605 of the information generating apparatus 600 may be further configured to: convert the above fused graph model into a corresponding vector matrix; determine the above fused graph model The vector corresponding to each node in the above, wherein the vector corresponding to each node above represents the characteristic information of the above parameter name or the above parameter content information; according to the above vector matrix and the vector corresponding to each above node, the summary information of the above target item is generated .
  • the second generating unit 605 of the information generating apparatus 600 may be further configured to: input the above-mentioned first output result into a pre-trained encoding network for processing time-series text data, to obtain The above second output result.
  • an electronic device 700 may include a processing device (eg, a central processing unit, a graphics processor, etc.) 701 that may be loaded into random access according to a program stored in a read only memory (ROM) 702 or from a storage device 707 Various appropriate actions and processes are executed by the programs in the memory (RAM) 703 . In the RAM 703, various programs and data required for the operation of the electronic device 700 are also stored.
  • the processing device 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704.
  • An input/output (I/O) interface 705 is also connected to bus 704 .
  • the processes described above with reference to the flowcharts may be implemented as computer software programs.
  • some embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the method illustrated in the flowchart.
  • the computer program may be downloaded and installed from a network via communication device 709, or from storage device 708, or from ROM 702.
  • the processing device 701 the above-mentioned functions defined in the methods of some embodiments of the present disclosure are performed.
  • a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
  • a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code therein.
  • the client and server can use any currently known or future developed network protocol such as HTTP (HyperText Transfer Protocol) to communicate, and can communicate with digital data in any form or medium Communication (eg, a communication network) interconnects.
  • HTTP HyperText Transfer Protocol
  • Examples of communication networks include local area networks (“LAN”), wide area networks (“WAN”), the Internet (eg, the Internet), and peer-to-peer networks (eg, ad hoc peer-to-peer networks), as well as any currently known or future development network of.
  • each block in the flowchart or block diagrams may represent a module, segment, or portion of code that contains one or more logical functions for implementing the specified functions executable instructions.
  • the functions noted in the blocks may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Computational Linguistics (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • General Health & Medical Sciences (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Mathematical Physics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Software Systems (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Evolutionary Computation (AREA)
  • Molecular Biology (AREA)
  • Computing Systems (AREA)
  • Probability & Statistics with Applications (AREA)
  • Library & Information Science (AREA)
  • Animal Behavior & Ethology (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

信息生成方法、装置、电子设备和计算机可读介质。该信息生成方法包括:获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据(401);对该序列化文本数据进行分词,得到词集(402);将该词集中目标词和该图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型(403);根据结构化文本数据和匹配后的图模型,生成融合后的图模型(404);根据该融合后的图模型,生成该目标物品的摘要信息(405)。该方法可以通过考量序列化文本数据和结构化文本数据的特征信息来准确、有效的生成目标物品的摘要信息。

Description

信息生成方法、装置、电子设备和计算机可读介质 技术领域
本公开的实施例涉及计算机技术领域,具体涉及信息生成方法、装置、电子设备和计算机可读介质。
背景技术
目前,传统的物品搜索与推荐的技术应用已经无法较好地满足用户日益增长的需求。在浏览推荐系统时,用户常常面临信息爆炸的问题。用户希望通过一些优致的物品摘要来快速了解产品。对于物品摘要的生成,通常采用的方式为:基于循环神经网络来生成物品的摘要。然而,循环神经网络只能建模序列文本信息,由于物品信息不仅包括非结构化的文本序列,也包括结构化信息。而循环神经网络对结构化输入难以直接建模,目前处理方法仍是作为字符串序列处理,这种处理方法可能会损失结构化信息,从而降低模型的表达能力。
发明内容
本公开的内容部分用于以简要的形式介绍构思,这些构思将在后面的具体实施方式部分被详细描述。本公开的内容部分并不旨在标识要求保护的技术方案的关键特征或必要特征,也不旨在用于限制所要求的保护的技术方案的范围。
本公开的一些实施例提出了信息生成方法、装置、设备和计算机可读介质,来解决以上背景技术部分提到的技术问题。
第一方面,本公开的一些实施例提供了一种信息生成方法,该方法包括:获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信 息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息;对上述序列化文本数据进行分词,得到词集;将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型;根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型;根据上述融合后的图模型,生成上述目标物品的摘要信息。
可选的,上述将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型,包括:将上述词集中目标词和上述图模型中的参数名称进行匹配,得到匹配结果;根据上述匹配结果和上述图模型中的目标边,生成匹配后的图模型,其中,上述目标边为上述图模型中表征上述目标词和上述目标词对应的参数名称之间关联程度信息的边。
可选的,上述根据上述融合后的图模型,生成上述目标物品的摘要信息,包括:将上述融合后的图模型转换为对应的向量矩阵;确定上述融合后的图模型中每个节点对应的向量,其中,上述每个节点对应的向量表征上述参数名称或上述参数内容信息的特征信息;根据上述向量矩阵和上述每个节点对应的向量,生成上述目标物品的摘要信息。
可选的,上述根据上述向量矩阵和上述每个节点对应的向量,生成上述目标物品的摘要信息,包括:将上述向量矩阵和上述融合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络,得到第一输出结果;对上述第一输出结果进行编码,得第二输出结果;对上述第二输出结果进行解码,得到上述目标物品的摘要信息。
可选的,上述对上述第一输出结果进行编码,得第二输出结果,包括:将上述第一输出结果输入至预先训练的、处理时序文本数据的编码网络,得到上述第二输出结果。
可选的,上述对上述第二输出结果进行解码,得到上述目标物品对应的摘要信息,包括:将上述第二输出结果输入至预先训练的、处理时序文本数据的解码网络,得到上述目标物品的摘要信息。
可选的,上述图模型中的边对应的数值通过以下步骤确定:获取 上述目标物品品类下的各个物品的物品信息,其中,上述物品信息包括参数名称和至少一个参数内容信息;根据上述各个物品的物品信息,确定每个参数名称对应的各个参数内容信息和每个参数内容信息在上述各个物品的物品信息中出现的次数;根据上述每个参数名称对应的各个参数内容信息和每个参数内容信息出现的频次,利用词频-逆向文件频率方法,确定上述图模型中的边对应的数值。
第二方面,本公开的一些实施例提供了一种信息生成装置,装置包括:获取单元,被配置成获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息;分词单元,被配置成对上述序列化文本数据进行分词,得到词集;匹配生成单元,被配置成将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型;第一生成单元,被配置成根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型;第二生成单元,被配置成根据上述融合后的图模型,生成上述目标物品的摘要信息。
可选的,匹配生成单元进一步配置成:将上述词集中目标词和上述图模型中的参数名称进行匹配,得到匹配结果;根据上述匹配结果和上述图模型中的目标边,生成匹配后的图模型,其中,上述目标边为上述图模型中表征上述目标词和上述目标词对应的参数名称之间关联程度信息的边。
可选的,第二生成单元进一步配置成:将上述融合后的图模型转换为对应的向量矩阵;确定上述融合后的图模型中每个节点对应的向量,其中,上述每个节点对应的向量表征上述参数名称或上述参数内容信息的特征信息;根据上述向量矩阵和上述每个节点对应的向量,生成上述目标物品的摘要信息。
可选的,第二生成单元进一步配置成:将上述向量矩阵和上述融 合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络,得到第一输出结果;对上述第一输出结果进行编码,得第二输出结果;对上述第二输出结果进行解码,得到上述目标物品的摘要信息。
可选的,第二生成单元进一步配置成:将上述第一输出结果输入至预先训练的、处理时序文本数据的编码网络,得到上述第二输出结果。
可选的,第二生成单元进一步配置成:将上述第二输出结果输入至预先训练的、处理时序文本数据的解码网络,得到上述目标物品的摘要信息。
第三方面,本公开的一些实施例提供了一种电子设备,包括:一个或多个处理器;存储装置,其上存储有一个或多个程序,当一个或多个程序被一个或多个处理器执行,使得一个或多个处理器实现如第一方面中任一的方法。
第四方面,本公开的一些实施例提供了一种计算机可读介质,其上存储有计算机程序,其中,程序被处理器执行时实现如第一方面中任一的方法。
本公开的上述各个实施例中具有如下有益效果:通过本公开的一些实施例的信息生成方法可以通过考量序列化文本数据和结构化文本数据的特征信息来准确、有效的生成目标物品的摘要信息。具体来说,循环神经网络只能建模序列文本信息,由于物品信息不仅包括非结构化的文本序列,也包括结构化信息。而循环神经网络对结构化输入难以直接建模,目前处理方法仍是作为字符串序列处理,这种处理方法可能会损失结构化信息,从而降低模型的表达能力。基于此,本公开的一些实施例的信息生成方法可以首先获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据。其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息。在这里,上述图模型用于确定后续词集中目标词对应的参数名称,为后续生成匹 配后的图模型奠定基础。需要说明的是,获取上述图模型用于将目标物品相关的序列化文本数据转化为结构化数据。然后,对上述序列化文本数据进行分词,得到词集。进而,将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。接着,根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。在这里,融合后的图模型体现了目标物品相关的结构化文本信息。最后,根据上述融合后的图模型,生成上述目标物品的摘要信息。由此可得,上述信息生成方法可以通过考量序列化文本数据和结构化文本数据的特征信息来准确、有效的生成目标物品的摘要信息。
附图说明
结合附图并参考以下具体实施方式,本公开各实施例的上述和其他特征、优点及方面将变得更加明显。贯穿附图中,相同或相似的附图标记表示相同或相似的元素。应当理解附图是示意性的,原件和元素不一定按照比例绘制。
图1-3是本公开的一些实施例的信息生成方法的一个应用场景图的示意图;
图4是根据本公开的信息生成方法一些实施例的流程图;
图5是根据本公开的信息生成方法的另一些实施例的流程图;
图6是根据本公开的信息生成装置的一些实施例的结构示意图;
图7是适于用来实现本公开的一些实施例的电子设备的结构示意图。
具体实施方式
下面将参照附图更详细地描述本公开的实施例。虽然附图中显示了本公开的某些实施例,然而应当理解的是,本公开可以通过各种形式来实现,而且不应该被解释为限于这里阐述的实施例。相反,提供这些实施例是为了更加透彻和完整地理解本公开。应当理解的是,本公开的附图及实施例仅用于示例性作用,并非用于限制本公开的保护范围。
另外还需要说明的是,为了便于描述,附图中仅示出了与有关发明相关的部分。在不冲突的情况下,本公开中的实施例及实施例中的特征可以相互组合。
需要注意,本公开中提及的“第一”、“第二”等概念仅用于对不同的装置、模块或单元进行区分,并非用于限定这些装置、模块或单元所执行的功能的顺序或者相互依存关系。
需要注意,本公开中提及的“一个”、“多个”的修饰是示意性而非限制性的,本领域技术人员应当理解,除非在上下文另有明确指出,否则应该理解为“一个或多个”。
本公开实施方式中的多个装置之间所交互的消息或者信息的名称仅用于说明性的目的,而并不是用于对这些消息或信息的范围进行限制。
下面将参考附图并结合实施例来详细说明本公开。
图1-图3是本公开的一些实施例的信息生成方法的一个应用场景图的示意图。
如图1-3所示,电子设备101可以首先获取预先构建的、与目标物品品类相关联的图模型105、目标物品102相关的序列化文本数据104和结构化文本数据103。其中,上述图模型105表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型105中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型105中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息。在本应用场景中,上述图模型105包括:“屏幕尺寸->6.5英寸,关联程度信息:0.6;屏幕尺寸->5.5英寸,关联程度信息:0.4;颜色->浅红色,关联程度信息:0.4;颜色->白色,关联程度信息:0.4;颜色->黑色,关联程度信息:0.3,电池容量->500,关联程度信息:0.3;电池容量->mHh,关联程度信息:0.9;电池容量->1000,关联程度信息:0.6;像素->500,关联程度信息:0.5;像素->1000,关联程度信息:0.4;像素->800,关联程度信息:0.5;像素->万,关联程度信息:0.8”。上述目标物品102可以是:“手机”。上述序列化文本数据104可以是:“手机品牌:**,颜色:红色系,电池 容量:3000mAh,像素:1000万”。结构化文本数据103可以包括:“**手机质量好,屏幕尺寸大,浅红色,电池容量大,防滑耐磨,像素高”。如图2所示,然后,对上述序列化文本数据104进行分词,得到词集106。进而,将上述词集106中目标词107和上述图模型105中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型108。在本应用场景中,词集106可以包括:“**手机”,“质量”,“好”,“屏幕尺寸”,“浅红色”“大”,“容量”,“大”,“防滑”,“耐磨”,“像素”,“高”。目标词107可以是:“浅红色”。上述匹配后的图模型108可以包括:“颜色->浅红色,关联程度信息:0.4;**手机;质量;好;屏幕尺寸;容量:大;耐磨:像素;高”。如图3所示,接着,根据上述结构化文本数据103和上述匹配后的图模型108,得到融合后的图模型109。最后,根据上述融合后的图模型109,生成上述目标物品102的摘要信息110。在本应用场景中,融合后的图模型109可以包括:“颜色->浅红色,关联程度信息:0.4;电池容量->1000,关联程度信息:0.6;像素->1000,关联程度信息:0.4;屏幕尺寸->5.5英寸,关联程度信息:0.4;**手机;质量;大;好;耐磨;高”。上述摘要信息110可以是:“**手机,质量好,屏幕尺寸大,5.5英寸,电池容量大,1000mAh,像素高,1000万”。
需要说明的是,上述电子设备101可以是硬件,也可以是软件。当电子设备为硬件时,可以实现成多个服务器或终端设备组成的分布式集群,也可以实现成单个服务器或单个终端设备。当电子设备体现为软件时,可以安装在上述所列举的硬件设备中。其可以实现成例如用来提供分布式服务的多个软件或软件模块,也可以实现成单个软件或软件模块。在此不做具体限定。
应该理解,图1中的电子设备的数目仅仅是示意性的。根据实现需要,可以具有任意数目的电子设备。
继续参考图4,示出了根据本公开的信息生成方法的一些实施例的流程400。该信息生成方法,包括以下步骤:
步骤401,获取预先构建的、与目标物品品类相关联的图模型、 目标物品相关的序列化文本数据和结构化文本数据。
在一些实施例中,信息生成方法的执行主体(例如图1所示的电子设备101)可以通过有线连接方式或者无线连接方式获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据。其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息。其中,上述关联程度信息可以是0-1之间的数值。
需要指出的是,上述无线连接方式可以包括但不限于3G/4G/5G连接、WiFi连接、蓝牙连接、WiMAX连接、Zigbee连接、UWB(ultra wideband)连接、以及其他现在已知或将来开发的无线连接方式。
步骤402,对上述序列化文本数据进行分词,得到词集。
在一些实施例中,上述执行主体可以对上述序列化文本数据进行分词,得到词集。作为示例,可以通过结巴分词来对上述序列化文本数据进行分词,得到词集。
步骤403,将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。
在一些实施例中,上述执行主体可以将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。其中,上述目标词可以是与图模型中的参数名称存在关联关系的词。作为示例,上述执行主体可以将匹配结果转换为图的形式,可以得到匹配后的图模型。
在一些实施例的一些可选的实现方式中,上述将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型,可以包括以下步骤:
第一步,将上述词集中目标词和上述图模型中的参数名称进行匹配,得到匹配结果。作为示例,上述执行主体可以首先提取上述词集中目标词的语义信息。然后,通过遍历图模型中的各个参数名称来确定与目标词的语义信息相关的参数名称,最后,将上述与目标词的语 义信息相关的参数名称与上述目标词进行匹配,得到匹配结果。
第二步,根据上述匹配结果和上述图模型中的目标边,生成匹配后的图模型。其中,上述目标边为上述图模型中表征上述目标词和上述目标词对应的参数名称之间关联程度信息的边。作为示例,上述执行主体可以根据上述匹配结果,确定上述匹配后的图模型的节点信息。最后,将上述图模型中的目标边作为上述匹配后的图模型的边,以此生成匹配后的图模型。
步骤404,根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。
在一些实施例的一些可选的实现方式中,上述执行主体可以根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。作为示例,上述执行主体可以将结构化文本数据融入上述匹配后的图模型,得到融合后的图模型。在这里,融合后的图模型体现了目标物品相关的结构化文本信息。
步骤405,根据上述融合后的图模型,生成上述目标物品的摘要信息。
在一些实施例中,上述执行主体可以根据上述融合后的图模型,生成上述目标物品的摘要信息。作为示例,上述执行主体可以首先将上述融合后的图模型转换为对应的向量。然后,将上述对应的向量和上述融合后的图模型中各个节点对应的向量输入至预先训练的循环神经网络,得到上述目标物品的摘要信息。
在一些实施例的一些可选的实现方式中,上述图模型中的边对应的数值通过以下步骤确定:
第一步,获取上述目标物品品类下的各个物品的物品信息,其中,上述物品信息包括参数名称和至少一个参数数值。
第二步,根据上述各个物品的物品信息,确定每个参数名称对应的各个参数内容信息和每个参数内容信息在上述各个物品的物品信息中出现的次数。在这里,各个物品的物品信息中每个物品信息对应着至少一个参数名称。进而,可以确定各个物品的物品信息对应参数名称集。其中,上述参数名称集中可能存在多个重复的参数名称。以此, 对参数名称集进行去重,可以得到去重后的参数名称集。在这里,去重后参数名称集中每个参数名称对应着至少一个参数内容信息。由此,通过各个物品的物品信息,可以确定每个参数内容信息在上述各个物品的物品信息中出现的次数。
第三步,根据上述每个参数名称对应的各个参数内容信息和每个参数内容信息出现的频次,利用词频-逆向文件频率方法,确定上述图模型中的边对应的数值。
在这里,词频-逆向文件频率方法的主要思想是,如果某个词或短语在一篇文章中出现的频率(Term Frequency,TF)高,并且在其他文章中很少出现,则认为此词或者短语具有很好的类别区分能力,适合用来分类。而逆向文件频率(Inverse Document Frequency,IDF)主要是指,如果包含某个词或短语的文档越少,则IDF越大,则说明该词或短语具有很好的类别区分能力。由此,使用词频-逆向文件频率方法,可以计算某个词或短语在某篇文章里面的重要性。
本公开的上述各个实施例中具有如下有益效果:通过本公开的一些实施例的信息生成方法可以通过考量序列化文本数据和结构化文本数据的特征信息来准确、有效的生成目标物品的摘要信息。具体来说,循环神经网络只能建模序列文本信息,由于物品信息不仅包括非结构化的文本序列,也包括结构化信息。而循环神经网络对结构化输入难以直接建模,目前处理方法仍是作为字符串序列处理,这种处理方法可能会损失结构化信息,从而降低模型的表达能力。基于此,本公开的一些实施例的信息生成方法可以首先获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据。其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息。在这里,上述图模型用于确定后续词集中目标词对应的参数名称,为后续生成匹配后的图模型奠定基础。需要说明的是,获取上述图模型用于将目标物品相关的序列化文本数据转化为结构化数据。然后,对上述序列化 文本数据进行分词,得到词集。进而,将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。接着,根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。在这里,融合后的图模型体现了目标物品相关的结构化文本信息。最后,根据上述融合后的图模型,生成上述目标物品的摘要信息。由此可得,上述信息生成方法可以通过考量序列化文本数据和结构化文本数据的特征信息来准确、有效的生成目标物品的摘要信息。
继续参考图5,示出了根据本公开的信息生成方法的另一些实施例的流程500。该信息生成方法,包括以下步骤:
步骤501,获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据。
步骤502,对上述序列化文本数据进行分词,得到词集。
步骤503,将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。
步骤504,根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。
在一些实施例中,步骤501-504的具体实现及所带来的技术效果可以参考图4对应的那些实施例中的步骤401-404,在此不再赘述。
步骤505,将上述融合后的图模型转换为对应的向量矩阵。
在一些实施例中,执行主体(例如图1所示的电子设备101)可以将上述融合后的图模型转换为对应的向量矩阵。作为示例,上述执行主体可以依据融合后的图模型中的各个边对应的数值来生成对应的向量矩阵。其中,对于孤立的节点与其他节点的关系可以表征为数值0。
步骤506,确定上述融合后的图模型中每个节点对应的向量。
在一些实施例中,上述执行主体可以确定上述融合后的图模型中每个节点对应的向量。其中,上述每个节点对应的向量表征上述参数名称或上述参数内容信息的特征信息。作为示例,可以通过词嵌入的方式来将上述融合后的图模型中每个节点对应的参数名称或参数内容 信息转换为对应的向量。
步骤507,将上述向量矩阵和上述融合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络,得到第一输出结果。
在一些实施例中,上述执行主体可以将上述向量矩阵和上述融合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络(GCN,Graph Convolutional Network),得到第一输出结果。其中,上述图卷积网络可以多层图卷积层。上述图卷积层可以通过邻接矩阵将每个节点的邻居信息进行聚合,进而增强当前节点的表征,提高模型对输入建模的准确性。进行多层图卷积层建模能获取多跳邻居的信息。
步骤508,对上述第一输出结果进行编码,得第二输出结果。
在一些实施例中,上述执行主体可以对上述第一输出结果进行编码,得第二输出结果。作为示例,上述执行主体可以将上述第一输出结果输入至预先训练的编码网络,得第二输出结果。
在一些实施例的一些可选的实现方式中,上述执行主体可以将上述第一输出结果输入至预先训练的、处理时序文本数据的编码网络,得到上述第二输出结果。其中,上述处理时序文本数据的编码网络可以是以下之一:循环神经网络(Recurrent Neural Network,RNN),长短期记忆网络(LSTM,Long Short-Term Memory)。
步骤509,对上述第二输出结果进行解码,得到上述目标物品的摘要信息。
在一些实施例中,上述执行主体可以对上述第二输出结果进行解码,得到上述目标物品的摘要信息。作为示例,上述执行主体可以将上述第二输出结果输入至预先训练的解码网络,得第二输出结果。
在一些实施例的一些可选的实现方式中,上述执行主体可以将上述第二输出结果输入至预先训练的、处理时序文本数据的解码网络,得到上述目标物品的摘要信息。其中,上述处理时序文本数据的解码网络可以是以下之一:循环神经网络,长短期记忆网络。
从图5中可以看出,与图4对应的一些实施例的描述相比,图5对应的一些实施例中的信息生成方法的流程500体现了根据上述融合 后的图模型,生成上述目标物品的摘要信息的步骤。由此,这些实施例描述的方案可以更精准、有效的生成上述目标物品的摘要信息。
继续参考图6,作为对上述各图上述方法的实现,本公开提供了一种信息生成装置的一些实施例,这些装置实施例与图4上述的那些方法实施例相对应,该装置具体可以应用于各种电子设备中。
如图6所示,一些实施例的信息生成装置600包括:获取单元601、分词单元602、匹配生成单元603、第一生成单元604和第二生成单元605。其中,获取单元601,被配置成获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,上述图模型表征上述目标物品品类的参数名称与各个参数数值之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息。分词单元602,被配置成对上述序列化文本数据进行分词,得到词集。匹配生成单元603,被配置成将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型。第一生成单元604,被配置成根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型。第二生成单元605,被配置成根据上述融合后的图模型,生成上述目标物品的摘要信息。
在一些实施例的一些可选的实现方式中,上述信息生成装置600的匹配生成单元603可以进一步被配置成:将上述词集中目标词和上述图模型中的参数名称进行匹配,得到匹配结果;根据上述匹配结果和上述图模型中的目标边,生成匹配后的图模型,其中,上述目标边为上述图模型中表征上述目标词和上述目标词对应的参数名称之间关联程度信息的边。
在一些实施例的一些可选的实现方式中,信息生成装置600的第二生成单元605可以进一步被配置成:将上述融合后的图模型转换为对应的向量矩阵;确定上述融合后的图模型中每个节点对应的向量,其中,上述每个节点对应的向量表征上述参数名称或上述参数内容信 息的特征信息;根据上述向量矩阵和上述每个节点对应的向量,生成上述目标物品的摘要信息。
在一些实施例的一些可选的实现方式中,信息生成装置600的第二生成单元605可以进一步被配置成:将上述向量矩阵和上述融合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络,得到第一输出结果;对上述第一输出结果进行编码,得第二输出结果;对上述第二输出结果进行解码,得到上述目标物品的摘要信息。
在一些实施例的一些可选的实现方式中,信息生成装置600的第二生成单元605可以进一步被配置成:将上述第一输出结果输入至预先训练的、处理时序文本数据的编码网络,得到上述第二输出结果。
在一些实施例的一些可选的实现方式中,信息生成装置600的第二生成单元605可以进一步被配置成:将上述第二输出结果输入至预先训练的、处理时序文本数据的解码网络,得到上述目标物品的摘要信息。
可以理解的是,该装置600中记载的诸单元与参考图4描述的方法中的各个步骤相对应。由此,上文针对方法描述的操作、特征以及产生的有益效果同样适用于装置600及其中包含的单元,在此不再赘述。
下面参考图7,其示出了适于用来实现本公开的一些实施例的电子设备700的结构示意图。图7示出的电子设备仅仅是一个示例,不应对本公开的实施例的功能和使用范围带来任何限制。
如图7所示,电子设备700可以包括处理装置(例如中央处理器、图形处理器等)701,其可以根据存储在只读存储器(ROM)702中的程序或者从存储装置707加载到随机访问存储器(RAM)703中的程序而执行各种适当的动作和处理。在RAM 703中,还存储有电子设备700操作所需的各种程序和数据。处理装置701、ROM 702以及RAM 703通过总线704彼此相连。输入/输出(I/O)接口705也连接至总线704。
通常,以下装置可以连接至I/O接口705:包括例如触摸屏、触摸 板、键盘、鼠标、摄像头、麦克风、加速度计、陀螺仪等的输入装置706;包括例如液晶显示器(LCD)、扬声器、振动器等的输出装置707;包括例如磁带、硬盘等的存储装置708;以及通信装置709。通信装置709可以允许电子设备700与其他设备进行无线或有线通信以交换数据。虽然图7示出了具有各种装置的电子设备700,但是应理解的是,并不要求实施或具备所有示出的装置。可以替代地实施或具备更多或更少的装置。图7中示出的每个方框可以代表一个装置,也可以根据需要代表多个装置。
特别地,根据本公开的一些实施例,上文参考流程图描述的过程可以被实现为计算机软件程序。例如,本公开的一些实施例包括一种计算机程序产品,其包括承载在计算机可读介质上的计算机程序,该计算机程序包含用于执行流程图所示的方法的程序代码。在这样的一些实施例中,该计算机程序可以通过通信装置709从网络上被下载和安装,或者从存储装置708被安装,或者从ROM 702被安装。在该计算机程序被处理装置701执行时,执行本公开的一些实施例的方法中限定的上述功能。
需要说明的是,本公开的一些实施例上述的计算机可读介质可以是计算机可读信号介质或者计算机可读存储介质或者是上述两者的任意组合。计算机可读存储介质例如可以是——但不限于——电、磁、光、电磁、红外线、或半导体的系统、装置或器件,或者任意以上的组合。计算机可读存储介质的更具体的例子可以包括但不限于:具有一个或多个导线的电连接、便携式计算机磁盘、硬盘、随机访问存储器(RAM)、只读存储器(ROM)、可擦式可编程只读存储器(EPROM或闪存)、光纤、便携式紧凑磁盘只读存储器(CD-ROM)、光存储器件、磁存储器件、或者上述的任意合适的组合。在本公开的一些实施例中,计算机可读存储介质可以是任何包含或存储程序的有形介质,该程序可以被指令执行系统、装置或者器件使用或者与其结合使用。而在本公开的一些实施例中,计算机可读信号介质可以包括在基带中或者作为载波一部分传播的数据信号,其中承载了计算机可读的程序代码。这种传播的数据信号可以采用多种形式,包括但不限于电磁信 号、光信号或上述的任意合适的组合。计算机可读信号介质还可以是计算机可读存储介质以外的任何计算机可读介质,该计算机可读信号介质可以发送、传播或者传输用于由指令执行系统、装置或者器件使用或者与其结合使用的程序。计算机可读介质上包含的程序代码可以用任何适当的介质传输,包括但不限于:电线、光缆、RF(射频)等等,或者上述的任意合适的组合。
在一些实施方式中,客户端、服务器可以利用诸如HTTP(HyperText Transfer Protocol,超文本传输协议)之类的任何当前已知或未来研发的网络协议进行通信,并且可以与任意形式或介质的数字数据通信(例如,通信网络)互连。通信网络的示例包括局域网(“LAN”),广域网(“WAN”),网际网(例如,互联网)以及端对端网络(例如,ad hoc端对端网络),以及任何当前已知或未来研发的网络。
上述计算机可读介质可以是上述装置中所包含的;也可以是单独存在,而未装配入该电子设备中。上述计算机可读介质承载有一个或者多个程序,当上述一个或者多个程序被该电子设备执行时,使得该电子设备:获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,上述图模型表征上述目标物品品类的参数名称与各个参数内容信息之间的关联关系,上述图模型中的节点表征上述目标物品品类的参数名称或参数内容信息,上述图模型中的边对应的数值表征上述参数名称与上述参数内容信息之间的关联程度信息;对上述序列化文本数据进行分词,得到词集;将上述词集中目标词和上述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型;根据上述结构化文本数据和上述匹配后的图模型,生成融合后的图模型;根据上述融合后的图模型,生成上述目标物品的摘要信息。
可以以一种或多种程序设计语言或其组合来编写用于执行本公开的一些实施例的操作的计算机程序代码,上述程序设计语言包括面向对象的程序设计语言—诸如Java、Smalltalk、C++,还包括常规的过程式程序设计语言—诸如“C”语言或类似的程序设计语言。程序代 码可以完全地在用户计算机上执行、部分地在用户计算机上执行、作为一个独立的软件包执行、部分在用户计算机上部分在远程计算机上执行、或者完全在远程计算机或服务器上执行。在涉及远程计算机的情形中,远程计算机可以通过任意种类的网络——包括局域网(LAN)或广域网(WAN)——连接到用户计算机,或者,可以连接到外部计算机(例如利用因特网服务提供商来通过因特网连接)。
附图中的流程图和框图,图示了按照本公开各种实施例的系统、方法和计算机程序产品的可能实现的体系架构、功能和操作。在这点上,流程图或框图中的每个方框可以代表一个模块、程序段、或代码的一部分,该模块、程序段、或代码的一部分包含一个或多个用于实现规定的逻辑功能的可执行指令。也应当注意,在有些作为替换的实现中,方框中所标注的功能也可以以不同于附图中所标注的顺序发生。例如,两个接连地表示的方框实际上可以基本并行地执行,它们有时也可以按相反的顺序执行,这依所涉及的功能而定。也要注意的是,框图和/或流程图中的每个方框、以及框图和/或流程图中的方框的组合,可以用执行规定的功能或操作的专用的基于硬件的系统来实现,或者可以用专用硬件与计算机指令的组合来实现。
描述于本公开的一些实施例中的单元可以通过软件的方式实现,也可以通过硬件的方式来实现。所描述的单元也可以设置在处理器中,例如,可以描述为:一种处理器包括:获取单元、分词单元、匹配生成单元、第一生成单元和第二生成单元。其中,这些单元的名称在某种情况下并不构成对该单元本身的限定,例如,分词单元还可以被描述为“对上述序列化文本数据进行分词,得到词集的单元”。
本文中以上描述的功能可以至少部分地由一个或多个硬件逻辑部件来执行。例如,非限制性地,可以使用的示范类型的硬件逻辑部件包括:现场可编程门阵列(FPGA)、专用集成电路(ASIC)、专用标准产品(ASSP)、片上系统(SOC)、复杂可编程逻辑设备(CPLD)等等。
以上描述仅为本公开的一些较佳实施例以及对所运用技术原理的说明。本领域技术人员应当理解,本公开的实施例中所涉及的发明范 围,并不限于上述技术特征的特定组合而成的技术方案,同时也应涵盖在不脱离上述发明构思的情况下,由上述技术特征或其等同特征进行任意组合而形成的其它技术方案。例如上述特征与本公开的实施例中公开的(但不限于)具有类似功能的技术特征进行互相替换而形成的技术方案。

Claims (10)

  1. 一种信息生成的方法,包括:
    获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,所述图模型表征所述目标物品品类的参数名称与各个参数内容信息之间的关联关系,所述图模型中的节点表征所述目标物品品类的参数名称或参数内容信息,所述图模型中的边对应的数值表征所述参数名称与所述参数内容信息之间的关联程度信息;
    对所述序列化文本数据进行分词,得到词集;
    将所述词集中目标词和所述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型;
    根据所述结构化文本数据和所述匹配后的图模型,生成融合后的图模型;
    根据所述融合后的图模型,生成所述目标物品的摘要信息。
  2. 根据权利要求1所述的方法,其中,所述将所述词集中目标词和所述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型,包括:
    将所述词集中目标词和所述图模型中的参数名称进行匹配,得到匹配结果;
    根据所述匹配结果和所述图模型中的目标边,生成匹配后的图模型,其中,所述目标边为所述图模型中表征所述目标词和所述目标词对应的参数名称之间关联程度信息的边。
  3. 根据权利要求1或2所述的方法,其中,所述根据所述融合后的图模型,生成所述目标物品的摘要信息,包括:
    将所述融合后的图模型转换为对应的向量矩阵;
    确定所述融合后的图模型中每个节点对应的向量,其中,所述每个节点对应的向量表征所述参数名称或所述参数内容信息的特征信 息;
    根据所述向量矩阵和所述每个节点对应的向量,生成所述目标物品的摘要信息。
  4. 根据权利要求3所述的方法,其中,所述根据所述向量矩阵和所述每个节点对应的向量,生成所述目标物品的摘要信息,包括:
    将所述向量矩阵和所述融合后的图模型中各个节点对应的向量输入至预先训练的图卷积神经网络,得到第一输出结果;
    对所述第一输出结果进行编码,得第二输出结果;
    对所述第二输出结果进行解码,得到所述目标物品的摘要信息。
  5. 根据权利要求4所述的方法,其中,所述对所述第一输出结果进行编码,得第二输出结果,包括:
    将所述第一输出结果输入至预先训练的、处理时序文本数据的编码网络,得到所述第二输出结果。
  6. 根据权利要求4或5所述的方法,其中,所述对所述第二输出结果进行解码,得到所述目标物品对应的摘要信息,包括:
    将所述第二输出结果输入至预先训练的、处理时序文本数据的解码网络,得到所述目标物品的摘要信息。
  7. 根据权利要求1-6之一所述的方法,其中,所述图模型中的边对应的数值通过以下步骤确定:
    获取所述目标物品品类下的各个物品的物品信息,其中,所述物品信息包括参数名称和至少一个参数数值;
    根据所述各个物品的物品信息,确定每个参数名称对应的各个参数内容信息和每个参数内容信息在所述各个物品的物品信息中出现的次数;
    根据所述每个参数名称对应的各个参数内容信息和每个参数内容信息出现的频次,利用词频-逆向文件频率方法,确定所述图模型中的 边对应的数值。
  8. 一种信息生成装置,包括:
    获取单元,被配置成获取预先构建的、与目标物品品类相关联的图模型、目标物品相关的序列化文本数据和结构化文本数据,其中,所述图模型表征所述目标物品品类的参数名称与各个参数数值之间的关联关系,所述图模型中的节点表征所述目标物品品类的参数名称或参数内容信息,所述图模型中的边对应的数值表征所述参数名称与所述参数内容信息之间的关联程度信息;
    分词单元,被配置成对所述序列化文本数据进行分词,得到词集;
    匹配生成单元,被配置成将所述词集中目标词和所述图模型中的参数名称进行匹配以及根据匹配结果生成匹配后的图模型;
    第一生成单元,被配置成根据所述结构化文本数据和所述匹配后的图模型,生成融合后的图模型;
    第二生成单元,被配置成根据所述融合后的图模型,生成所述目标物品的摘要信息。
  9. 一种电子设备,包括:
    一个或多个处理器;
    存储装置,用于存储一个或多个程序;
    当所述一个或多个程序被所述一个或多个处理器执行时,使得所述一个或多个处理器实现如权利要求1-7中任一所述的方法。
  10. 一种计算机可读介质,其上存储有计算机程序,其中,所述程序被处理器执行时实现如权利要求1-7中任一所述的方法。
PCT/CN2021/140219 2021-02-19 2021-12-21 信息生成方法、装置、电子设备和计算机可读介质 Ceased WO2022174669A1 (zh)

Priority Applications (2)

Application Number Priority Date Filing Date Title
JP2023550238A JP7656057B2 (ja) 2021-02-19 2021-12-21 情報生成方法、装置、電子機器及びコンピュータ可読媒体
US18/277,928 US12306858B2 (en) 2021-02-19 2021-12-21 Information generation method, apparatus, electronic device, and computer-readable medium

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202110188846.2 2021-02-19
CN202110188846.2A CN113779316B (zh) 2021-02-19 2021-02-19 信息生成方法、装置、电子设备和计算机可读介质

Publications (1)

Publication Number Publication Date
WO2022174669A1 true WO2022174669A1 (zh) 2022-08-25

Family

ID=78835609

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2021/140219 Ceased WO2022174669A1 (zh) 2021-02-19 2021-12-21 信息生成方法、装置、电子设备和计算机可读介质

Country Status (4)

Country Link
US (1) US12306858B2 (zh)
JP (1) JP7656057B2 (zh)
CN (1) CN113779316B (zh)
WO (1) WO2022174669A1 (zh)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116416056A (zh) * 2023-04-04 2023-07-11 深圳征信服务有限公司 一种基于机器学习的征信数据处理方法及系统
CN117743293A (zh) * 2023-11-14 2024-03-22 国网物资有限公司 数据存储方法、装置、电子设备和计算机可读介质

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113779316B (zh) * 2021-02-19 2026-03-20 北京沃东天骏信息技术有限公司 信息生成方法、装置、电子设备和计算机可读介质
CN116881438B (zh) * 2022-03-28 2026-04-17 北京沃东天骏信息技术有限公司 生成客服咨询摘要的方法和装置

Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10146751B1 (en) * 2014-12-31 2018-12-04 Guangsheng Zhang Methods for information extraction, search, and structured representation of text data
CN111563117A (zh) * 2020-07-14 2020-08-21 北京每日优鲜电子商务有限公司 结构化信息显示方法、装置、电子设备和计算机可读介质
CN111723215A (zh) * 2020-06-19 2020-09-29 国家计算机网络与信息安全管理中心 基于文本挖掘的生物技术信息知识图谱构建装置与方法
CN111914157A (zh) * 2019-05-07 2020-11-10 重庆好德译信息技术有限公司 一种基于用户偏好的旅游方案生成方法和系统
CN113779316A (zh) * 2021-02-19 2021-12-10 北京沃东天骏信息技术有限公司 信息生成方法、装置、电子设备和计算机可读介质

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8438178B2 (en) * 2008-06-26 2013-05-07 Collarity Inc. Interactions among online digital identities
US9857960B1 (en) * 2015-08-25 2018-01-02 Palantir Technologies, Inc. Data collaboration between different entities
US10909313B2 (en) * 2016-06-22 2021-02-02 Sas Institute Inc. Personalized summary generation of data visualizations
WO2020003174A2 (en) * 2018-06-29 2020-01-02 3M Innovative Properties Company Semantic graph textual coding
CN111046135A (zh) * 2019-11-04 2020-04-21 智器云南京信息科技有限公司 非结构文本处理方法、装置、计算机设备、存储介质
CN111080399A (zh) * 2019-11-22 2020-04-28 汉口北进出口服务有限公司 一种商品信息处理方法和装置
CN111523304B (zh) * 2020-04-27 2022-08-02 华东师范大学 一种基于预训练模型的产品描述文本的自动生成方法
US20230186120A1 (en) * 2021-11-24 2023-06-15 Intelligent Fusion Technology, Inc. Methods and systems for anomaly and pattern detection of unstructured big data

Patent Citations (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10146751B1 (en) * 2014-12-31 2018-12-04 Guangsheng Zhang Methods for information extraction, search, and structured representation of text data
CN111914157A (zh) * 2019-05-07 2020-11-10 重庆好德译信息技术有限公司 一种基于用户偏好的旅游方案生成方法和系统
CN111723215A (zh) * 2020-06-19 2020-09-29 国家计算机网络与信息安全管理中心 基于文本挖掘的生物技术信息知识图谱构建装置与方法
CN111563117A (zh) * 2020-07-14 2020-08-21 北京每日优鲜电子商务有限公司 结构化信息显示方法、装置、电子设备和计算机可读介质
CN113779316A (zh) * 2021-02-19 2021-12-10 北京沃东天骏信息技术有限公司 信息生成方法、装置、电子设备和计算机可读介质

Cited By (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN116416056A (zh) * 2023-04-04 2023-07-11 深圳征信服务有限公司 一种基于机器学习的征信数据处理方法及系统
CN116416056B (zh) * 2023-04-04 2023-10-03 深圳征信服务有限公司 一种基于机器学习的征信数据处理方法及系统
CN117743293A (zh) * 2023-11-14 2024-03-22 国网物资有限公司 数据存储方法、装置、电子设备和计算机可读介质

Also Published As

Publication number Publication date
JP7656057B2 (ja) 2025-04-02
CN113779316A (zh) 2021-12-10
US12306858B2 (en) 2025-05-20
CN113779316B (zh) 2026-03-20
US20240232237A9 (en) 2024-07-11
JP2024509077A (ja) 2024-02-29
US20240134892A1 (en) 2024-04-25

Similar Documents

Publication Publication Date Title
WO2022174669A1 (zh) 信息生成方法、装置、电子设备和计算机可读介质
CN111104599B (zh) 用于输出信息的方法和装置
CN119150862B (zh) 模型微调方法、文本处理方法、介质、设备及程序产品
CN120277206B (zh) 基于知识图谱的检索增强生成方法、装置、设备及介质
CN111090993A (zh) 属性对齐模型训练方法及装置
WO2022156730A1 (zh) 文本处理方法、装置、设备及介质
CN115700548A (zh) 用户行为预测的方法、设备和计算机程序产品
CN113807926B (zh) 推荐信息生成方法、装置、电子设备和计算机可读介质
CN116128055A (zh) 图谱构建方法、装置、电子设备和计算机可读介质
CN118520115A (zh) 基于rpa和ai技术的信息管理系统及方法
CN113946648B (zh) 结构化信息生成方法、装置、电子设备和介质
CN111626044B (zh) 文本生成方法、装置、电子设备及计算机可读存储介质
CN113780516A (zh) 物品文案生成网络训练方法、物品文案生成方法、装置
WO2022017299A1 (zh) 一种文本检测方法、装置、电子设备及存储介质
CN115481260B (zh) 基于审计信息的知识图谱构建方法、装置和电子设备
CN118171303A (zh) 模型数据存储方法、装置、电子设备和计算机可读介质
CN119648818A (zh) 图像生成方法、装置、系统、电子设备、计算机可读介质
CN112148865A (zh) 信息推送方法和装置
CN119671271B (zh) 价值数据处理方法、装置、电子设备和计算机可读介质
CN118113703B (zh) 数据处理方法、装置、电子设备和计算机可读介质
CN119693941B (zh) 食物描述信息显示方法、装置、设备和计算机可读介质
CN111860518B (zh) 用于分割图像的方法、装置、设备和计算机可读介质
CN113742479A (zh) 筛选目标文本的方法和装置
CN120407755B (zh) 请求应答方法、装置、电子设备以及存储介质
CN113779314B (zh) 信息处理方法、装置、电子设备和计算机可读介质

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 21926382

Country of ref document: EP

Kind code of ref document: A1

WWE Wipo information: entry into national phase

Ref document number: 18277928

Country of ref document: US

WWE Wipo information: entry into national phase

Ref document number: 2023550238

Country of ref document: JP

NENP Non-entry into the national phase

Ref country code: DE

WWE Wipo information: entry into national phase

Ref document number: 11202306184P

Country of ref document: SG

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 24.11.2023)

122 Ep: pct application non-entry in european phase

Ref document number: 21926382

Country of ref document: EP

Kind code of ref document: A1

WWG Wipo information: grant in national office

Ref document number: 18277928

Country of ref document: US