WO2019196228A1 - 制度信息处理方法、装置、计算机设备和存储介质 - Google Patents

制度信息处理方法、装置、计算机设备和存储介质 Download PDF

Info

Publication number
WO2019196228A1
WO2019196228A1 PCT/CN2018/095549 CN2018095549W WO2019196228A1 WO 2019196228 A1 WO2019196228 A1 WO 2019196228A1 CN 2018095549 W CN2018095549 W CN 2018095549W WO 2019196228 A1 WO2019196228 A1 WO 2019196228A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
system information
extended
target
category
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/CN2018/095549
Other languages
English (en)
French (fr)
Inventor
韩梅
张安元
邓华威
王科
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Ping An Technology Shenzhen Co Ltd
Original Assignee
Ping An Technology Shenzhen Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Ping An Technology Shenzhen Co Ltd filed Critical Ping An Technology Shenzhen Co Ltd
Publication of WO2019196228A1 publication Critical patent/WO2019196228A1/zh
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/06Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/289Phrasal analysis, e.g. finite state techniques or chunking
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/237Lexical tools
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06QINFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES; SYSTEMS OR METHODS SPECIALLY ADAPTED FOR ADMINISTRATIVE, COMMERCIAL, FINANCIAL, MANAGERIAL OR SUPERVISORY PURPOSES, NOT OTHERWISE PROVIDED FOR
    • G06Q10/00Administration; Management
    • G06Q10/06Resources, workflows, human or project management; Enterprise or organisation planning; Enterprise or organisation modelling
    • G06Q10/063Operations research, analysis or management
    • G06Q10/0639Performance analysis of employees; Performance analysis of enterprise or organisation operations

Definitions

  • the present application relates to a system information processing method, apparatus, computer device and storage medium.
  • Institutional norms are the rules and guidelines that employees must abide by in their production and business activities, including laws and policies, organizational structure, management systems, job responsibilities, technical standards, and work processes.
  • institutions In order to meet the job requirements of each position, enterprises need to classify and manage the system from different dimensions, and build a number of different information trees, such as technical standard information trees, legal policy information trees, etc., so that different types and uses of different systems are formed. system. When a new system is released, the newly released system needs to be included in the corresponding information tree.
  • the same system may belong to multiple different information trees at the same time.
  • the corresponding institutional information and information trees are increasing.
  • the inventors realized that in the traditional way, it is not only inefficient and error-prone to manually manage the massive amount of institutional information.
  • a system information processing method, apparatus, computer device, and storage medium are provided.
  • An institutional information processing method includes: monitoring system information issued by a terminal, and classifying the system information to obtain a corresponding original word set; the original word set includes a plurality of original words; synthesizing each original word synonymously, generating a set of extended words corresponding to each original word; forming an extended system information set corresponding to the system information according to each extended word set; inputting the extended system information set into a preset system management model to obtain a target corresponding to the system information And selecting a category label corresponding to each of the plurality of target information trees, filtering a target information tree including the category label corresponding to the target category, and adding the system information to the filtered target information tree.
  • An institutional information processing device includes:
  • An information expansion module configured to monitor system information issued by the terminal, and perform word segmentation on the system information to obtain a corresponding original word set;
  • the original word set includes a plurality of original words; synonymously expand each original word to generate each a set of extended words corresponding to the original word; forming a set of extended system information corresponding to the system information according to each set of extended words;
  • An information classification module configured to input the extended system information set into a preset system management model, to obtain a target category corresponding to the system information;
  • the information archiving module is configured to obtain a category label corresponding to each of the plurality of target information trees, filter a target information tree including the category label corresponding to the target category, and add the system information to the filtered target information tree.
  • a computer apparatus comprising a memory and one or more processors having stored therein computer readable instructions that, when executed by a processor, implement the steps of the system information processing method provided in any one of the embodiments of the present application.
  • One or more non-volatile storage media storing computer readable instructions, when executed by one or more processors, causing one or more processors to implement a system as provided in any one embodiment of the present application The steps of the information processing method.
  • FIG. 1 is an application scenario diagram of a system information processing method according to one or more embodiments.
  • FIG. 2 is a flow diagram of a method of processing institutional information in accordance with one or more embodiments.
  • FIG. 3 is a schematic diagram of a target information tree in a system information processing method in accordance with one or more embodiments.
  • FIG. 4 is a flow diagram of the steps of constructing an associated information tree in accordance with one or more embodiments.
  • FIG. 5 is a schematic diagram of an associated information tree in a system information processing method in accordance with one or more embodiments.
  • FIG. 6 is a block diagram showing the structure of an institutional information processing apparatus according to one or more embodiments.
  • FIG. 7 is a block diagram of a computer device in accordance with one or more embodiments.
  • the system information processing method provided by the present application can be applied to an application environment as shown in FIG. 1.
  • the terminal 102 communicates with the server 104 over a network.
  • the terminal 102 can be, but is not limited to, various personal computers, notebook computers, smart phones, tablets, and portable wearable devices, and the server 104 can be implemented with a stand-alone server or a server cluster composed of a plurality of servers.
  • a plurality of target information trees are stored in the server 104. Each target information tree has a corresponding category label.
  • the server 104 monitors whether the terminal 102 issues new system information. When it is detected that the terminal 102 issues new system information, the server 104 classifies the system information and incorporates the system information into the corresponding target information tree. Specifically, the server 104 segments the system information to obtain a collection of original words including a plurality of original words. The server 104 obtains synonym corresponding to each original word, and forms an expanded word set with the original word and the corresponding synonym. There is a corresponding set of extended words for each original word.
  • the server 104 arbitrarily selects one word from the set of extended words corresponding to each original word in the order in which the original words appear in the system information, and forms an extended system information in order. When different words are selected from the expanded word set, different extended system information is formed, and different extended system information constitutes an expanded system information set.
  • the server 104 inputs the expanded system information set into the trained system management model, and uses the system management model to determine the target category corresponding to the system information.
  • the server 104 acquires the category label corresponding to the target category, filters the information node including the obtained category label, and adds the system information to the filtered target information tree.
  • the segmentation and synonym expansion of institutional information improves the effective coverage of institutional information, so that after inputting the system management model, the accuracy of the target category can be improved, and the system information can be accurately incorporated into the corresponding target information tree, and the system can be improved.
  • Information classification efficiency and accuracy improves the effective coverage of institutional information, so that after inputting the system management model, the accuracy of the target category can be improved, and the system information can be accurately incorporated into the corresponding target information tree, and the system can be improved.
  • a system information processing method is provided.
  • the method is applied to the server in FIG. 1 as an example, and includes the following steps:
  • Step 202 Monitor system information issued by the terminal, and classify the system information to obtain a corresponding original word set; the original word set includes a plurality of original words.
  • the server monitors whether the first terminal issues new system information.
  • Institutional information includes institutional description information and associated institutional documents.
  • the system description information includes the system code, the system name, the system level, the issuing unit, the release date, the applicable object identifier or the information summary.
  • the system information may be text information, voice information, image information, video information, and the like. If it is voice information, image information or video information, voice information, image information and video information can be converted into text information by voice recognition or image processing.
  • the institutional document includes a number of institutional provisions and the applicable object identifier for each system clause. Applicable object identification refers to the identification information of the object that needs to perform or understand the system, and may be a post identification or an organization identification.
  • the server classifies the system information. Specifically, the server classifies the system information by using a word segmentation algorithm to obtain a collection of original words.
  • the original set of words includes a plurality of original words.
  • words that have a small effect on the classification such as stop words, modal particles, and punctuation marks, are removed, thereby improving the efficiency of subsequent feature extraction.
  • a stop word refers to a word in the system information that appears more than a preset threshold but has little practical meaning, such as me, he, etc.
  • the method before the system information is segmented to obtain the corresponding original word set, the method further includes: detecting whether the system description information includes category information; if included, adding the system information to the corresponding target information tree according to the category information. Otherwise, the system information is segmented to obtain the corresponding original word set.
  • the terminal may also pre-specify the category information of the system information, so that the server can incorporate the system information into the corresponding target information tree according to the category information. If the system description information does not include the category information of the system information, the system information may be classified and managed according to the system information processing method provided by the present application.
  • Step 204 Synonymously expand each original word to generate an extended word set corresponding to each original word.
  • the server separately obtains the synonym corresponding to each original word in the original word set, and forms the extended word set by the original word and the corresponding synonym.
  • Synonyms refer to words that have the same or similar meaning as the original words.
  • the original words are “not allowed”, and the synonyms can be “no”, “forbidden”, “avoided”, “cancelled”, etc., and the original words and corresponding synonyms are formed. Expand the collection of words, such as the set of extended words corresponding to the original word "not allowed” as ⁇ not, no, prohibited, avoided, put an end to ⁇ .
  • each original word in the original word set has a corresponding extended word set, such as a corresponding extended word set of a is ⁇ a, a1, a2 ⁇ , b corresponding
  • the set of extended words is ⁇ b, b1, b2, b3 ⁇
  • the set of extended words corresponding to c is ⁇ c, c1, c2 ⁇ .
  • Step 206 Form an extended system information set corresponding to the system information according to each extended word set.
  • the server arbitrarily selects one word from the set of extended words corresponding to each original word according to the order in which the original words appear in the system information, and forms an extended system information in order.
  • different words are selected from the expanded word set, different extended system information is formed, and different extended system information constitutes an expanded system information set.
  • the server obtains a Cartesian product for the set of extended words corresponding to each original word, and forms a set of extended system information composed of different extended system information.
  • the Cartesian product of the two sets X and Y also known as the direct product, is expressed as X ⁇ Y.
  • the first object is a member of X and the second object is one of all possible ordered pairs of Y.
  • Step 208 Input the extended system information set into a preset system management model, and obtain a target category corresponding to the system information.
  • the institutional management model is for determining a target category corresponding to the input from among a plurality of candidate types based on the input.
  • the system management model can be a model obtained by training such as logistic regression algorithm and support vector machine algorithm.
  • the system management model can be formed by multiple sub-management model connections. Since the input of the trained system management model is an expanded set of extended system information, the expanded information of each extended system expresses the same or similar meaning as the institutional information, and improves the effective coverage of the institutional information, so as to be input later. After the trained system management model, the accuracy of the target category can be improved.
  • Step 210 Acquire a category label corresponding to each of the plurality of target information trees, filter the target information tree including the category label corresponding to the target category, and add the system information to the target information tree obtained by the screening.
  • each target information tree includes a plurality of information nodes and an institutional file associated with each information node.
  • Institutional files can be files in multiple formats, such as pdf documents, jpg images, xls tables, mp3 audio or avi videos.
  • Different information nodes can be arranged in the target information tree according to the release time. It is easy to understand that an institutional information may also have no associated institutional documents, and may also have multiple associated institutional documents, which is not limited.
  • Each target information tree has a corresponding category label.
  • the category label is used to identify the category of information nodes that the corresponding target information tree can contain, such as administrative management, sales management, or risk management.
  • the server obtains the category annotation corresponding to the target category, and filters one or more target information trees including the obtained category annotation.
  • the server generates an information node based on the system description information. For example, the system number and/or the system name can be used as information nodes.
  • the server associates the system file to the information node, and adds the information node associated with the system file to the target information tree obtained by the screening.
  • the institutional information includes an institutional description information and an institutional document; adding the institutional information to the filtered target information tree includes: generating an information node according to the institutional description information; and detecting whether the same target information tree in the screening has been Information node; if not present, the information node is added to the corresponding target information tree, and the system file is associated with the information node.
  • the server determines, according to the system description information, whether the generated information node belongs to a parallel node or a parent child node with the same information node that already exists. When the generated information node and the existing information node belong to the parallel node, the server distinguishes the generated information node from the existing information node, and adds the marked information node to the corresponding target information tree.
  • the system file is associated with the information node after the difference mark.
  • the server When the generated information node and the existing same information node belong to the parallel node, the server describes and defines the generated information node according to the system description information, that is, extracts the keyword in the system description information, and generates the generated keyword pair.
  • the information node performs semantic expansion.
  • the information node generated according to the name of the system is the “company welfare management system”
  • the keyword “research and development department” is extracted from the system description information
  • the information node after the semantic expansion may be the “company development department welfare management system”.
  • the server adds the semantically expanded information node as a child node of the existing same information node to the corresponding target information tree, and associates the system file to the child node.
  • the system information is segmented to obtain the corresponding original word set; by obtaining the synonym corresponding to each original word in the original word set, the original word and the corresponding synonym can be used to form the extended word set.
  • the extended system information set corresponding to the institutional information can be formed; the extended system information set is input into the trained system management model to obtain the target category corresponding to the institutional information; and the target category and the pre-stored are more
  • the target information trees are matched by the corresponding category labels, and the target information tree capable of including the system information can be filtered, and the system information is added to the filtered target information tree.
  • the extended word set corresponding to each original word is formed, and then the expanded system information set is formed by expanding the word set, which greatly improves the expansion of the extended system information.
  • the expanded extended system information expresses the same or similar meaning as the institutional information. Improve the effective coverage of institutional information, so that after the input of the trained system management model, the accuracy of the target category can be improved, and the system information can be accurately incorporated into the corresponding target information tree, and the efficiency and accuracy of the system information classification can be improved. rate.
  • the generating step of the system management model includes: acquiring training sample data; the training sample data includes a plurality of sample system information and corresponding category labels; performing segmentation and synonymous expansion processing on each sample system information to obtain Each sample system information corresponds to the extended sample system information set; according to each extended sample system information set and corresponding category labeling, the initial system management model is trained by the support vector machine algorithm to obtain the system management model.
  • the training sample data can be published sample system information.
  • Each sample system information has a corresponding category label that describes the actual category of sample system information. For example, if the name of the system corresponding to the sample system information is “ attendance considerations”, the category label corresponding to the sample system information may be “administrative management”.
  • the training sample data includes sample system information corresponding to all possible categories to ensure the accuracy of each category determination. In a specific embodiment, the training sample data includes 476 sample system information, and the total number of category annotations is 57.
  • the server uses the word segmentation algorithm to segment each training sample information to obtain each word, and each word constitutes a set of original training words corresponding to each training sample information.
  • the server obtains a synonym of each original training word, and forms an extended training word set with the original training word and the corresponding synonym.
  • Extended training word set includes multiple groups
  • the server first obtains one training sample information as current training sample information, acquires each original training word corresponding to the current training sample information, obtains an extended training word set corresponding to each original training word, and then presses each original training word in the current training sample information.
  • a word is arbitrarily selected from the set of extended training words corresponding to each original training word, and an extended sample system information is formed in order.
  • Different extended sample system information constitutes an expanded sample system information set.
  • Each sample system information has a corresponding set of extended sample system information.
  • the server obtains a Cartesian product for the set of extended training words corresponding to each original training word, and forms an extended sample system information set corresponding to each sample system information.
  • the support vector machine algorithm is a machine learning algorithm used for pattern recognition and pattern classification.
  • the main idea of the support vector machine is to establish an optimal decision hyperplane, so that the distance between the two types of samples on the two sides of the plane from the plane is maximized, thus providing a good generalization ability for the classification problem.
  • the system randomly generates a hyperplane and moves continuously to classify the samples until the sample points belonging to different categories in the training sample are located on both sides of the hyperplane. There may be many hyperplanes that satisfy the condition.
  • the support vector machine algorithm finds such a hyperplane while ensuring the classification accuracy, so that the white space on both sides of the hyperplane is maximized, thus achieving the optimal classification of linear separable samples.
  • the support vector machine algorithm is a kind. Supervised training methods.
  • the institutional management model is formed by a plurality of sub-management model connections.
  • a large number of published system information is segmented and synonymously expanded, and the obtained extended sample system information set is processed, which greatly improves the effective coverage of the sample system information; and the expanded sample system information set is input into the system management model. And based on the support vector machine algorithm to train the system management model, the classification accuracy of the system management model can be improved.
  • the extended sample system information set includes a plurality of sets of extended sample system information; and the initial system management model is trained by the support vector machine algorithm according to each extended sample system information set and corresponding category labeling: acquiring features Item, calculating the word frequency weight of the feature item in a set of extended sample system information; calculating the document frequency of the feature item in the entire training sample data; calculating the feature weight corresponding to the feature item according to the word frequency weight and the document frequency; selecting the feature item according to the feature weight as Correspondingly expanding the feature words of the sample system information; extracting the characteristics of each extended sample standard information according to the feature words.
  • the feature item may be any one of a set of extended sample system information.
  • the word frequency weight refers to the frequency at which feature items appear in the group of extended sample system information. It can be understood that if there is a synonym of the feature item in the extended sample system information, it also appears.
  • the word frequency weight is usually normalized and can be expressed as TF ij , where i represents the identity corresponding to the feature item and j represents the class identifier.
  • the document frequency DF i is a measure of the universal importance of a word, which can be obtained by dividing the number of extended sample system information in which the feature item is located by the total number of all training sample information in the training sample data.
  • the feature weight is proportional to the word frequency weight. If the number of extended sample system information appearing in the feature item is more, it indicates that the effect of the feature item on the information classification is smaller, that is, the feature weight is inversely proportional to the document frequency.
  • the feature weight Where N represents the total number of all training sample information in the training sample data.
  • the feature weight exceeds the preset threshold, it indicates that the feature item is an important term of the extended sample system information, and the feature item can be used as the feature word of the extended sample system information.
  • the feature of each extended sample system information in the extended sample system information set may be extracted according to the determined feature words.
  • the feature words may include one or more.
  • the institutional information includes an institutional description information and an associated institutional document; the institutional document includes a plurality of institutional terms and corresponding applicable object identifiers; and the associated information tree has a corresponding applicable object identifier.
  • the method also includes the step of constructing an associated information tree. As shown in FIG. 4, the steps of constructing the associated information tree include:
  • step 402 the system file is split, and the system sub-file corresponding to the applicable object identifier is generated by using the system clause corresponding to each applicable object identifier.
  • the company may record all the system information applicable to different positions to the same system file, so that users can only conduct system inquiry based on the entire information content of the system documents, thereby reducing the efficiency of system information query.
  • This embodiment constructs different association information trees for different positions.
  • the server splits the multiple system terms in the system file according to the applicable object identifier corresponding to each system clause in the system file, and generates a system sub-file corresponding to each applicable object identifier.
  • the institutional document A includes four system clauses X1 to X4.
  • X1 corresponds to the applicable object identifier including A and B
  • X2 corresponds to the applicable object identifier including A
  • X3 corresponds to the applicable object identifier including A, B, C, D and E
  • X4 corresponds to the applicable object identifier including A and D.
  • Institutional Document A consists of five applicable object identifiers: A, B, C, D and E. The corresponding splits are obtained in five system sub-documents A1 to A5.
  • the system sub-file A1 corresponding to the applicable object identifier A includes four system clauses X1 to X4;
  • the system sub-file A2 corresponding to the applicable object identifier B includes two system clauses of X1 and X3; and so on.
  • Step 404 Acquire multiple association information trees corresponding to the target information tree.
  • Each target information tree has a corresponding plurality of associated information trees.
  • Each information node in the target information tree has a corresponding one or more applicable object identifiers.
  • the different applicable object identifiers in the target information tree respectively have a corresponding associated information tree.
  • the number of applicable object identifiers in the target information tree is equal to the number of corresponding association information trees, so that each applicable object identifier corresponding post has a corresponding associated information tree.
  • the target information tree is used to record institutional information that applies to all positions in the enterprise.
  • the associated information tree only needs to record the institutional information applicable to a position.
  • Each associated information tree has a corresponding applicable object identifier.
  • the associated information tree corresponding to the object identifier "post 1" is applied, and the information node 4 and the information are not present in the target information tree of FIG. Node 9. It is easy to understand that the directory hierarchy of multiple information nodes in the associated information tree is not necessarily consistent with the target information tree, and can be adaptively adjusted. The content of the system file record associated with other information nodes still existing in the associated information tree may be different from the content of the system file record associated with the corresponding information node in the target information tree.
  • Step 406 Add the system description information and the system sub-file to the corresponding association information tree according to the applicable object identifier.
  • the server After the server adds the system information to the corresponding target information tree, the server obtains the corresponding associated information tree corresponding to the target information tree according to the applicable object identifier recorded by the system file. It is easy to understand that the server only needs to obtain the associated information tree corresponding to the applicable object identifier of the system file record.
  • the institutional information classification is added to three target information trees, including the target information tree M.
  • the target information tree M corresponds to the applicable object identifiers including A, B, C, D, E, and E.
  • the system file only includes information content applicable to A, B, C, D, and E according to the above example, the server only needs to obtain the target information.
  • the associated information tree corresponding to A, B, C, D, and E corresponding to tree M.
  • the server generates an information node according to the system description information, and associates the split multiple system sub-files with the information node.
  • the server adds multiple information nodes associated with different system sub-files to the associated information tree corresponding to the same applicable object identifier. For example, in the above example, the association has added system subfile information of the nodes A1 to object information tree M applicable object ID A corresponding association information tree M A; and is associated with adding system subfile information of the node A2 to the target information
  • the tree M corresponds to the associated information tree M B corresponding to the object identifier B , and so on.
  • the server When receiving the system query request sent by the second terminal, the server acquires the associated information tree corresponding to the applicable object identifier.
  • the system query request carries the applicable object identifier and query conditions.
  • the server searches for an information node that satisfies the query condition in the association information tree, acquires a system sub-file associated with the information node that satisfies the query condition, and sends the system sub-file to the second terminal.
  • the system documents that are applicable to the system information of different positions are recorded, and the system clauses that need to be executed or understood for each position are selected to meet the individual needs of different positions.
  • Different posts are used to construct an associated information tree that only contains the content of the corresponding post requirements, and the process of generating all associated information trees is fully automated, saving time and effort; subsequent users only need to perform system query based on the associated information tree applicable to them, and can also improve Institutional query efficiency.
  • splitting the system file includes: calculating a data amount of the system file, detecting whether the data amount exceeds a threshold; when the data amount exceeds the threshold, acquiring a preset target data amount, and determining a system according to the target data amount The split position of the file; detects whether the split position is between adjacent separators; when the split position is at a separator, splits the system file into multiple intermediate files at the split position; when the split position is located When the adjacent separators are separated, the system file is split into multiple intermediate files at any one of the adjacent separators; multiple intermediate files are split according to the preset splitting rules.
  • the server calculates the amount of data in the system file and checks whether the amount of data exceeds the threshold.
  • the threshold may be preset or may be temporarily generated based on the load monitoring result of the server.
  • the server may pre-separate the system file into multiple intermediate files with a small amount of data, and then split the intermediate files into multiple system sub-files.
  • the server acquires a preset target data amount, and determines a split location of the system file according to the target data amount.
  • the target data amount may be preset or may be temporarily generated based on load monitoring results of other servers in the plurality of clusters. For example, the data volume of the system file A is 720M. If the target data volume is 80M, the 80M size position of the system file is marked as the first split position, and the 160M size position is marked as the second split position. And so on.
  • the server identifies if each split location is between adjacent separators. When the split location is located at a location where the separator is located, the server splits the system file at the split location to obtain a plurality of intermediate files corresponding to the system file. When the split position is between adjacent separators, the server splits the corresponding system file at any one of the adjacent separators, that is, the previous separator or the next separator in the adjacent separator The symbol is split to obtain multiple intermediate files corresponding to the system file.
  • the server calls multi-threading to split the intermediate file into multiple system sub-files as described above, or send the intermediate files to other servers in the cluster for splitting to improve file splitting efficiency.
  • the system file with a large amount of data is split into intermediate files with a small amount of data and then transmitted to other servers in the cluster for splitting, which can also improve data transmission efficiency.
  • two-level splitting is performed on the system file with a large amount of data: wherein the split of the first level is split according to the amount of data, and the split of the second level is performed according to the preset split dimension.
  • Splitting splitting; splitting the system file with a large amount of data into intermediate files with a small amount of data, and splitting the intermediate file into multiple system sub-files in parallel, thereby improving the efficiency of file splitting.
  • an institutional information processing apparatus including: an information expansion module 602, an information classification module 604, and an information archiving module 606, wherein:
  • the information expansion module 602 is configured to monitor system information issued by the terminal, and perform word segmentation on the system information to obtain a corresponding original word set; the original word set includes a plurality of original words; synonymously expand each original word to generate each original word corresponding a set of extended words; a set of extended system information corresponding to the system information is formed according to each set of extended words.
  • the information classification module 604 is configured to input the extended system information set into a preset system management model, and obtain a target category corresponding to the system information.
  • the information archiving module 606 is configured to obtain a category label corresponding to each of the plurality of target information trees, filter the target information tree including the category label corresponding to the target category, and add the system information to the filtered target information tree.
  • the system information includes the system description information; the information expansion module 602 is further configured to detect whether the system description information includes category information; if included, the system information is added to the corresponding target information tree according to the category information; otherwise, The system information is segmented to obtain the corresponding original word set.
  • the apparatus further includes a model training module 608, configured to acquire training sample data; the training sample data includes a plurality of sample system information and corresponding category labels; and segmentation and synonym expansion of each sample system information Processing, obtaining the extended sample system information set corresponding to each sample system information; according to each extended sample system information set and corresponding category labeling, the initial system management model is trained by the support vector machine algorithm to obtain the system management model.
  • a model training module 608 configured to acquire training sample data; the training sample data includes a plurality of sample system information and corresponding category labels; and segmentation and synonym expansion of each sample system information Processing, obtaining the extended sample system information set corresponding to each sample system information; according to each extended sample system information set and corresponding category labeling, the initial system management model is trained by the support vector machine algorithm to obtain the system management model.
  • the extended sample system information set includes a plurality of sets of extended sample system information; the model training module 608 is further configured to acquire feature items, calculate word frequency weights of the feature items in a set of extended sample system information; calculate feature items throughout The document frequency in the sample data is trained; the feature weight corresponding to the feature item is calculated according to the word frequency weight and the document frequency; the feature item is selected as the feature word of the corresponding extended sample system information according to the feature weight; and the feature of each extended sample standard information is extracted according to the feature word.
  • the system information includes the system description information and the system file; the information archiving module 606 is further configured to generate an information node according to the system description information; and detect whether the same information node exists in the target information tree obtained by the screening; If present, the information node is added to the corresponding target information tree, and the system file is associated to the information node.
  • the system information includes the system description information and the associated system file; the system file includes a plurality of system terms and corresponding applicable object identifiers; the associated information tree has a corresponding applicable object identifier; the information archiving module 606 also uses The system file is split, and the system sub-file corresponding to the applicable object identifier is generated by using the system clause corresponding to each applicable object identifier; the plurality of associated information trees corresponding to the target information tree are obtained; and the system description information is obtained according to the applicable object identifier. And the system sub-file is added to the corresponding associated information tree.
  • the information archiving module 606 is further configured to calculate a data amount of the system file, and detect whether the data amount exceeds a threshold; when the data amount exceeds the threshold, acquire a preset target data amount, and determine the system file according to the target data amount.
  • Split position detect whether the split position is between adjacent separators; when the split position is at a separator, split the system file into multiple intermediate files at the split position; when the split position is at the phase
  • the adjacent separators are separated, the system file is split into multiple intermediate files at any one of the adjacent separators; multiple intermediate files are split according to the preset splitting rules.
  • system information processing device For the specific definition of the system information processing device, reference may be made to the above limitation of the system information processing method, and details are not described herein again.
  • system information processing apparatuses may be implemented in whole or in part by software, hardware, and a combination thereof.
  • Each of the above modules may be embedded in or independent of the processor in the computer device, or may be stored in a memory in the computer device in a software form, so that the processor invokes the operations corresponding to the above modules.
  • a computer device which may be a server, and its internal structure diagram may be as shown in FIG.
  • the computer device includes a processor, memory, network interface, and database connected by a system bus.
  • the processor of the computer device is used to provide computing and control capabilities.
  • the memory of the computer device includes a non-volatile storage medium, an internal memory.
  • the non-volatile storage medium stores an operating system, computer readable instructions, and a database.
  • the internal memory provides an environment for operation of an operating system and computer readable instructions in a non-volatile storage medium.
  • the database of the computer device is used to store institutional information.
  • the network interface of the computer device is used to communicate with an external terminal via a network connection. The steps of the system information processing method provided in any one of the embodiments of the present application when the computer readable instructions are executed by the processor.
  • FIG. 7 is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation of the computer device to which the solution of the present application is applied.
  • the specific computer device may It includes more or fewer components than those shown in the figures, or some components are combined, or have different component arrangements.
  • One or more non-volatile storage media storing computer readable instructions, when executed by one or more processors, causing one or more processors to implement a system as provided in any one embodiment of the present application The steps of the information processing method.
  • Non-volatile memory can include read only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory.
  • Volatile memory can include random access memory (RAM) or external cache memory.
  • RAM is available in a variety of formats, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronization chain.
  • SRAM static RAM
  • DRAM dynamic RAM
  • SDRAM synchronous DRAM
  • DDRSDRAM double data rate SDRAM
  • ESDRAM enhanced SDRAM
  • Synchlink DRAM SLDRAM
  • Memory Bus Radbus
  • RDRAM Direct RAM
  • DRAM Direct Memory Bus Dynamic RAM
  • RDRAM Memory Bus Dynamic RAM

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Business, Economics & Management (AREA)
  • Human Resources & Organizations (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Strategic Management (AREA)
  • Entrepreneurship & Innovation (AREA)
  • Economics (AREA)
  • Educational Administration (AREA)
  • Artificial Intelligence (AREA)
  • Development Economics (AREA)
  • Health & Medical Sciences (AREA)
  • General Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • General Engineering & Computer Science (AREA)
  • Game Theory and Decision Science (AREA)
  • Marketing (AREA)
  • Operations Research (AREA)
  • Quality & Reliability (AREA)
  • Tourism & Hospitality (AREA)
  • General Business, Economics & Management (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种制度信息处理方法,包括:监测终端发布的制度信息,对制度信息进行分词得到对应的原始词语集合;原始词语集合包括多个原始词语;对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;根据各个扩展词语集合形成制度信息对应的扩展制度信息集合;将扩展制度信息集合输入预设的制度管理模型,得到制度信息对应的目标类别;获取多个目标信息树分别对应的类别标注,筛选包含与目标类别对应类别标注的目标信息树,将制度信息添加至筛选得到的目标信息树。

Description

制度信息处理方法、装置、计算机设备和存储介质
本申请要求于2018年4月9日提交中国专利局,申请号为201810313040X,申请名称为“制度信息处理方法、装置、计算机设备和存储介质”的中国专利申请的优先权,其全部内容通过引用结合在本申请中。
技术领域
本申请涉及一种制度信息处理方法、装置、计算机设备和存储介质。
背景技术
企业标准化是对企业生产经营与管理等活动中的重复性事物和概念,通过制订、发布和实施制度规范达到统一,以提高企业管理水平。制度规范(以下简称“制度”)是员工在生产经营活动中须共同遵守的规定和准则,包括法律与政策、企业组织结构、管理制度、岗位职责、技术标准、工作流程等规范文件。为了满足各个岗位的工作需求,企业需要从不同维度对制度进行分类管理,构建多个不同的信息树,如技术标准信息树、法律政策信息树等,以使不同类和用途的制度形成不同的制度体系。当有新制度发布时,需要将新发布的制度纳入相应的信息树。同一种制度可能同时隶属于多个不同的信息树。随着企业规模增大,相应的制度信息和信息树均越来越多。但发明人意识到,在传统的方式中,通过人工对海量制度信息进行分类管理,不仅效率低,且易出错。
发明内容
根据本申请公开的各种实施例,提供一种制度信息处理方法、装置、计算机设备和存储介质。
一种制度信息处理方法包括:监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;及获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标 注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
一种制度信息处理装置包括:
信息扩展模块,用于监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;
信息分类模块,用于将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;及
信息归档模块,用于获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
一种计算机设备,包括存储器和一个或多个处理器,存储器中存储有计算机可读指令,计算机可读指令被处理器执行时实现本申请任意一个实施例中提供的制度信息处理方法的步骤。
一个或多个存储有计算机可读指令的非易失性存储介质,计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器实现本申请任意一个实施例中提供的制度信息处理方法的步骤。
本申请的一个或多个实施例的细节在下面的附图和描述中提出。本申请的其它特征和优点将从说明书、附图以及权利要求书变得明显。
附图说明
为了更清楚地说明本申请实施例中的技术方案,下面将对实施例中所需要使用的附图作简单地介绍,显而易见地,下面描述中的附图仅仅是本申请的一些实施例,对于本领域普通技术人员来讲,在不付出创造性劳动的前提下,还可以根据这些附图获得其它的附图。
图1为根据一个或多个实施例中制度信息处理方法的应用场景图。
图2为根据一个或多个实施例中制度信息处理方法的流程示意图。
图3为根据一个或多个实施例中制度信息处理方法中目标信息树的示意图。
图4为根据一个或多个实施例中构建关联信息树步骤的流程示意图。
图5为根据一个或多个实施例中制度信息处理方法中关联信息树的示意图。
图6为根据一个或多个实施例中制度信息处理装置的结构框图。
图7为根据一个或多个实施例中计算机设备的框图。
具体实施方式
为了使本申请的技术方案及优点更加清楚明白,以下结合附图及实施例,对本申请进行进一步详细说明。应当理解,此处描述的具体实施例仅仅用以解释本申请,并不用于限定本申请。
本申请提供的制度信息处理方法,可以应用于如图1所示的应用环境中。终端102与服务器104通过网络进行通信。终端102可以但不限于是各种个人计算机、笔记本电脑、智能手机、平板电脑和便携式可穿戴设备,服务器104可以用独立的服务器或者是多个服务器组成的服务器集群来实现。
服务器104中存储了多种目标信息树。每种目标信息树具有对应的类别标注。服务器104对终端102是否发布新的制度信息进行监测,当监测到终端102发布了新的制度信息时,服务器104对制度信息进行分类,将制度信息纳入相应的目标信息树。具体的,服务器104对制度信息进行分词,得到包括多个原始词语的原始词语集合。服务器104获取各个原始词语对应的同义词,将原始词语与对应的同义词形成扩展词语集合。每个原始词语都存在对应的扩展词语集合。服务器104按与制度信息中各个原始词语出现的顺序,从各个原始词语对应的扩展词语集合中任意选择一个词语,按顺序形成一个扩展制度信息。当从扩展词语集合中选择不同的词语时,则形成不同的扩展制度信息,不同的扩展制度信息组成扩展制度信息集合。服务器104将扩展制度信息集合输入已训练的制度管理模型,利用制度管理模型确定制度信息对应的目标类别。服务器104获取与目标类别对应的类别标注,筛选包含获取到的类别标注的信息节点,将制度信息添加至筛选得到的目标信息树。对制度信息进行分词和同义扩展,提高了制度信息的有效覆盖范围,从而在后续输入制度管理模型后,可提高目标类别的精准性,进而可准确将制度信息纳入相应目标信息树,提高制度信息分类效率和准确率。
在其中一个实施例中,如图2所示,提供了一种制度信息处理方法,以该方法应用于图1中的服务器为例进行说明,包括以下步骤:
步骤202,监测终端发布的制度信息,对制度信息进行分词得到对应的原始 词语集合;原始词语集合包括多个原始词语。
服务器对第一终端是否发布新的制度信息进行监测。制度信息包括制度描述信息及关联的制度文件。制度描述信息包括制度编码、制度名称、制度级别、发布单位、发布日期、适用对象标识或信息摘要等。制度信息可以是文本信息,也可以是语音信息、图像信息、视频信息等。如果是语音信息、图像信息或视频信息,则可先通过语音识别或图像处理,将语音信息、图像信息和视频信息转化为文本信息。制度文件包括多项制度条款以及每项制度条款对应的适用对象标识。适用对象标识是指需要执行或了解该制度的对象的标识信息,可以是岗位标识或机构标识等。
当监测到第一终端发布了新的制度信息时,服务器对制度信息进行分类。具体的,服务器通过分词算法对制度信息进行分词,得到原始词语集合。原始词语集合包括多个原始词语。在其中一个实施例中,得到各个原始词语后,去除停用词、语气词、标点符号等对分类影响作用小的词语,从而提高后续特征提取的效率。停用词指的是制度信息中出现频率超过预设阈值但实际意义不大的词,如我,的,他等。
在其中一个实施例中,在对制度信息进行分词得到对应的原始词语集合之前,还包括:检测制度描述信息是否包含类别信息;若包含,则根据类别信息将制度信息添加至相应的目标信息树;否则,对制度信息进行分词得到对应的原始词语集合。
终端在发布制度信息时,也可以预先标明制度信息的类别信息,以便服务器可以根据该类别信息,将制度信息纳入相应的目标信息树。若制度描述信息并未包含制度信息的类别信息,则可以按照本申请提供的制度信息处理方法对制度信息进行分类管理。
步骤204,对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合。
服务器分别获取原始词语集合中各个原始词语对应的同义词,将原始词语与对应的同义词形成扩展词语集合。每个原始词语都存在对应的扩展词语集合。同义词是指与原始词语含义相同或相近的词语,如原始词语为“不得”,同义词可为“切勿”、“禁止”、“避免”、“杜绝”等,将原始词语与对应的同义词形成扩展词语集合,如原始词语“不得”对应的扩展词语集合为{不得,切勿,禁止, 避免,杜绝}。如原始词语集合为{a,b,c},则原始词语集合中的每个原始词语都存在对应的扩展词语集合,如a对应的扩展词语集合为{a,a1,a2},b对应的扩展词语集合为{b,b1,b2,b3},c对应的扩展词语集合为{c,c1,c2}。
步骤206,根据各个扩展词语集合形成制度信息对应的扩展制度信息集合。
服务器按照与制度信息中各个原始词语出现的顺序,从各个原始词语对应的扩展词语集合中任意选择一个词语,按顺序形成一个扩展制度信息。当从扩展词语集合中选择不同的词语时,则形成不同的扩展制度信息,不同的扩展制度信息组成扩展制度信息集合。
在其中一个实施例中,服务器对各个原始词语对应的扩展词语集合求笛卡尔积,形成由不同的扩展制度信息组成的扩展制度信息集合。两个集合X和Y的笛卡尔积,又称直积,表示为X×Y。第一个对象是X的成员而第二个对象是Y的所有可能有序对的其中一个成员。
步骤208,将扩展制度信息集合输入预设的制度管理模型,得到制度信息对应的目标类别。
制度管理模型用于根据输入从多个候选类型中确定与输入对应的目标类别。制度管理模型可以是通过逻辑回归算法、支持向量机算法等训练得到的模型。制度管理模型内部可以由多个子管理模型连接形成。由于已训练的制度管理模型的输入是经过扩展了的扩展制度信息集合,扩展后的各个扩展制度信息表达了与制度信息相同或相近的含义,提高了制度信息的有效覆盖范围,从而在后续输入已训练的制度管理模型后,可提高目标类别的精准性。
步骤210,获取多个目标信息树分别对应的类别标注,筛选包含与目标类别对应类别标注的目标信息树,将制度信息添加至筛选得到的目标信息树。
服务器中存储了多种目标信息树。如图3所示,每种目标信息树包括多个信息节点及每个信息节点关联的制度文件。制度文件可以是多种格式的文件,如pdf文档、jpg图像、xls表格、mp3音频或avi视频等。不同的信息节点在目标信息树中可以按照发布时间先后排列。容易理解,一项制度信息也可以不具有关联的制度文件,也还可以具有多个关联的制度文件,对此不作限制。
每种目标信息树具有对应的类别标注。类别标注用于标识相应目标信息树能够包含的信息节点的类别,如行政管理类、销售管理类或风险管理类等。服务器获取与目标类别对应的类别标注,筛选包含获取到的类别标注的一种或多 种目标信息树。服务器根据制度描述信息生成信息节点。例如,可以将制度编号和/或制度名称作为信息节点。服务器将制度文件关联至该信息节点,将关联有制度文件的信息节点添加至筛选得到的目标信息树。
在其中一个实施例中,制度信息包括制度描述信息和制度文件;将制度信息添加至筛选得到的目标信息树包括:根据制度描述信息生成信息节点;检测筛选得到的目标信息树中是否已存在相同的信息节点;若不存在,则将信息节点添加至相应的目标信息树,将制度文件关联至信息节点。
若筛选得到的关联信息树中已经存在相应的信息节点,则服务器只需将制度文件关联至已存在相应的信息节点。在另一个实施例中,服务器根据制度描述信息判断生成的信息节点与已存在的相同信息节点属于并列节点还是父子节点。当生成的信息节点与已存在的相同信息节点属于并列节点时,服务器对生成的信息节点与已存在的相同信息节点进行区别标记,将区别标记后的信息节点添加至相应的目标信息树,将制度文件关联至区别标记后的信息节点。
当生成的信息节点与已存在的相同信息节点属于并列节点时,服务器根据制度描述信息对生成的信息节点进行描述限定,即在制度描述信息中提取关键词,利用提取到的关键词对生成的信息节点进行语义扩充。例如,根据制度名称生成的信息节点为“公司福利管理制度”,在制度描述信息中提取关键词“研发部”,则语义扩充后的信息节点可以是“公司研发部福利管理制度”。服务器将语义扩充后的信息节点作为已存在的相同信息节点的子节点添加至相应的目标信息树,将制度文件关联至该子节点。
本实施例中,通过监测新发布的制度信息,对制度信息进行分词得到对应的原始词语集合;通过获取原始词语集合中各个原始词语对应的同义词,可以利用原始词语与对应的同义词形成扩展词语集合;根据各个原始词语对应的扩展词语集合,可以形成制度信息对应的扩展制度信息集合;将扩展制度信息集合输入已训练的制度管理模型得到制度信息对应的目标类别;将目标类别与预存储的多个目标信息树分别对应的类别标注进行匹配,可以筛选得到能够包含该制度信息的目标信息树,将制度信息添加至筛选得到的目标信息树。先形成每个原始词语对应的扩展词语集合,再通过扩展词语集合形成扩展制度信息集合,大大提高了扩展制度信息的扩展度,扩展后的各个扩展制度信息表达了与制度信息相同或相近的含义,提高了制度信息的有效覆盖范围,从而在后续输 入已训练的制度管理模型后,可提高目标类别的精准性,进而可以准确将制度信息纳入相应的目标信息树,提高制度信息分类效率和准确率。
在其中一个实施例中,制度管理模型的生成步骤包括:获取训练样本数据;训练样本数据包括多个样本制度信息及分别对应的类别标注;对各个样本制度信息进行分词和同义扩展处理,得到每个样本制度信息分别对应的扩展样本制度信息集合;根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,得到制度管理模型。
训练样本数据可以是已发布的多种样本制度信息。每种样本制度信息都有对应的类别标注,用于描述样本制度信息的实际类别。例如,样本制度信息对应的制度名称为“考勤注意事项”,则该样本制度信息对应的类别标注可以是“行政管理”。训练样本数据包括所有可能的类别对应的样本制度信息,以保证各个类别确定的准确性。在一个具体的实施例中,训练样本数据包括476个样本制度信息,类别标注总数为57。
服务器通过分词算法对各个训练样本信息进行分词得到各个词语,各个词语组成各个训练样本信息对应的原始训练词语集合。服务器获取每个原始训练词语的同义词,将原始训练词语与对应的同义词形成扩展训练词语集合。扩展训练词语集合包括多组
服务器先获取其中一个训练样本信息作为当前训练样本信息,获取当前训练样本信息对应的各个原始训练词语,获取各个原始训练词语对应的扩展训练词语集合,然后按与当前训练样本信息中各个原始训练词语出现的顺序,从各个原始训练词语对应的扩展训练词语集合中任意选择一个词语,按顺序形成一个扩展样本制度信息。不同的扩展样本制度信息组成扩展样本制度信息集合。各个样本制度信息都有对应的扩展样本制度信息集合。在其中一个实施例中,服务器对各个原始训练词语对应的扩展训练词语集合求笛卡尔积,形成得到每个样本制度信息分别对应的扩展样本制度信息集。
支持向量机算法是一种用来进行模式识别,模式分类的机器学习算法。支持向量机的主要思想是:建立一个最优决策超平面,使得该平面两侧距离该平面最近的两类样本之间的距离最大化,从而对分类问题提供良好的泛化能力。对于一个多维的样本集,系统随机产生一个超平面并不断移动,对样本进行分类,直到训练样本中属于不同类别的样本点正好位于该超平面的两侧,满足该 条件的超平面可能有很多个,支持向量机算法在保证分类精度的同时,寻找到这样一个超平面,使得超平面两侧的空白区域最大化,从而实现对线性可分样本的最优分类,支持向量机算法是一种有监督的训练方法。在其中一个实施例中,制度管理模型由多个子管理模型连接形成。
本实施例中,对大量已发布制度信息进行分词和同义扩展处理,处理得到的扩展样本制度信息集合,大大提高了样本制度信息的有效覆盖范围;将扩展样本制度信息集合输入制度管理模型,并基于支持向量机算法对制度管理模型进行训练,可提高制度管理模型的分类精准性。
在其中一个实施例中,扩展样本制度信息集合包括多组扩展样本制度信息;根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练包括:获取特征项,计算特征项在一组扩展样本制度信息的词频权重;计算特征项在整个训练样本数据中的文档频率;根据词频权重和文档频率计算特征项对应的特征权重;根据特征权重选择特征项作为相应扩展样本制度信息的特征词;根据特征词提取各个扩展样本标准信息的特征。
特征项可以是一组扩展样本制度信息中的任一个词语。词频权重指的是特征项在该组扩展样本制度信息中出现的频率。可以理解的是,如果扩展样本制度信息中存在特征项的同义词,则也算出现。词频权重通常被归一化,可表示为TF ij,其中i表示特征项对应的标识,j表示类别标识。文档频率DF i是一个词语普遍重要性的度量,可以由特征项所在的扩展样本制度信息数目除以训练样本数据中所有训练样本信息的总数目得到。
如果该特征项在扩展样本制度信息中出现的次数越多,表明该特征项对扩展样本制度信息的影响力度越大,即特征权重与词频权重成正比。如果该特征项出现的扩展样本制度信息的数量越多,表明,该特征项对信息分类的作用越小,即特征权重与文档频率成反比。在其中一个实施例中,特征权重
Figure PCTCN2018095549-appb-000001
其中N表示训练样本数据中所有训练样本信息的总数目。
如果特征权重超过预设阈值,则说明此特征项是这一组扩展样本制度信息的重要词语,可将此特征项作为此扩展样本制度信息的特征词。可根据确定的各个特征词提取扩展样本制度信息集合中各个扩展样本制度信息的特征。对于一个扩展样本制度信息,特征词可以包括一个或多个。
本实施例中,通过统计扩展样本制度信息中每个词语的词频权重和文档频 率,确定该词语代表扩展样本制度信息的特征权重,根据特征权重可以提取扩展样本制度信息集合中每个扩展样本制度信息的特征,从而使制度管理模型能够从基于多样化的语言描述的制度信息中准确提取其特征,进而进行准确分类。
在其中一个实施例中,制度信息包括制度描述信息及关联的制度文件;制度文件包括多个制度条款以及分别对应的适用对象标识;关联信息树具有对应的适用对象标识。该方法还包括构建关联信息树的步骤。如图4所示,构建关联信息树的步骤包括:
步骤402,对制度文件进行拆分,利用每个适用对象标识对应的制度条款生成相应适用对象标识对应的制度子文件。
为了满足所有岗位的工作需求,企业可能将适用于不同岗位的制度信息全部记录至同一个制度文件中,使得用户只能基于制度文件全部信息内容进行制度查询,进而使制度信息查询效率降低。本实施例针对不同岗位构建不同的关联信息树。具体的,服务器根据制度文件中每个制度条款对应的适用对象标识,对制度文件中多个制度条款进行拆分,生成每个适用对象标识分别对应的制度子文件。例如,制度文件A包括X1~X4四项制度条款。其中,X1对应适用对象标识包括甲和乙,X2对应适用对象标识包括甲,X3对应适用对象标识包括甲、乙、丙、丁和戊,X4对应适用对象标识包括甲和丁。制度文件A共包括甲、乙、丙、丁和戊五个适用对象标识,对应的拆分得到五个制度子文件A1~A5。其中,适用对象标识甲对应的制度子文件A1包括X1~X4四项制度条款;适用对象标识乙对应的制度子文件A2包括X1和X3两项制度条款;如此类推。
步骤404,获取目标信息树对应的多个关联信息树。
每种目标信息树具有对应的多个关联信息树。目标信息树中每个信息节点具有对应的一个或多个适用对象标识。目标信息树中不同适用对象标识分别具有对应的一个关联信息树。换言之,目标信息树中包含适用对象标识的数量与对应的关联信息树的数量相等,从而每个适用对象标识对应岗位具有对应的关联信息树。
目标信息树用于记录适用于企业全部岗位的制度信息。而关联信息树则只需记录适用于一个岗位的制度信息。每种关联信息树具有对应的适用对象标识。如图5所示,岗位1无需执行或了解信息节点4和信息节点9对应的制度,则适用对象标识“岗位1”对应的关联信息树,相对图3目标信息树不存在信息节 点4和信息节点9。容易理解,关联信息树中多个信息节点的目录层级,并非一定与目标信息树一致,可以自适应调整。关联信息树仍存在的其他信息节点关联的制度文件记录的内容,与目标信息树中相应信息节点关联的制度文件记录的内容可以不同。
步骤406,根据适用对象标识,将制度描述信息及制度子文件添加至相应的关联信息树。
服务器将制度信息添加至相应的目标信息树后,服务器根据制度文件记录的适用对象标识,获取目标信息树对应的相应关联信息树。容易理解,服务器只需获取制度文件记录的适用对象标识对应的关联信息树。例如,制度信息分类添加至三种目标信息树,其中包括目标信息树M。目标信息树M对应适用对象标识包括甲、乙、丙、丁、戊和己,假设依上述举例制度文件只包括适用于甲、乙、丙、丁和戊的信息内容,则服务器只需获取目标信息树M对应的甲、乙、丙、丁和戊分别对应的关联信息树。
服务器根据制度描述信息生成信息节点,将拆分得到的多个制度子文件分别关联至信息节点。服务器将多个关联有不同制度子文件的信息节点分别添加至相同适用对象标识对应的关联信息树。例如,在上述举例中,将关联有制度子文件A1的信息节点添加至目标信息树M中适用对象标识甲对应的关联信息树M ;将关联有制度子文件A2的信息节点添加至目标信息树M中适用对象标识乙对应的关联信息树M ,如此类推。
当接收到第二终端发送的制度查询请求时,服务器获取适用对象标识对应的关联信息树。制度查询请求携带了适用对象标识和查询条件。服务器在关联信息树中查找满足查询条件的信息节点,获取与满足查询条件的信息节点关联的制度子文件,将制度子文件发送至第二终端。
本实施例中,在制度信息发布时,将记录来了适用于不同岗位的制度信息的制度文件拆分,将每个岗位需要执行或了解的制度条款挑选出来,满足不同岗位个性化需求,为不同岗位分别构建只包含相应岗位需求内容的关联信息树,且所有关联信息树的生成过程全自动进行,省时省力;后续用户只需基于适用于自己的关联信息树进行制度查询,也可以提高制度查询效率。
在其中一个实施例中,对制度文件进行拆分包括:计算制度文件的数据量,检测数据量是否超过阈值;当数据量超过阈值时,获取预设的目标数据量,根 据目标数据量确定制度文件的拆分位置;检测拆分位置是否位于相邻分隔符之间;当拆分位置位于一个分隔符处时,在拆分位置将制度文件拆分为多个中间文件;当拆分位置位于相邻分隔符之间时,在相邻分隔符中任意一个分隔符处将制度文件拆分为多个中间文件;按照预设的拆分规则,对多个中间文件进行拆分。
服务器计算制度文件的数据量,检测数据量是否超过阈值。该阈值可以是预先设定的,也可以是根据服务器的负载监测结果临时生成的。当数据量超过阈值时,服务器可以将制度文件预先拆分为多个数据量小的中间文件,再将中间文件分别拆分为多个制度子文件。具体的,服务器获取预设的目标数据量,根据目标数据量确定制度文件的拆分位置。目标数据量可以是预先设定的,也可以是根据对多个集群内其他服务器的负载监测结果临时生成的。例如,制度文件A的数据量为720M,假设目标数据量为80M,则将制度文件的第80M大小的位置标记为第一个拆分位置,第160M大小的位置标记为第二个拆分位置,以此类推。
服务器识别每个拆分位置是否位于相邻分隔符之间。当拆分位置位于一个分隔符所在的位置时,服务器在该拆分位置对制度文件进行拆分,得到该制度文件对应的多个中间文件。当拆分位置位于相邻分隔符之间时,服务器在相邻分隔符中任意一个分隔符处对相应制度文件进行拆分,即对该相邻分隔符中的前一个分隔符或后一个分隔符处进行拆分,得到制度文件对应的多个中间文件。服务器调用多线程按照上述方式将中间文件拆分为多个制度子文件,或者将中间文件发送至集群内其他服务器进行拆分,以提高文件拆分效率。将数据量较大的制度文件拆分为数据量较小的中间文件后传输至集群内其他服务器进行拆分,还可以提高数据传输效率,
本实施例中,对于数据量较大的制度文件进行两级拆分:其中,第一层级的拆分是根据数据量进行拆分,第二层级的拆分是根据预设的拆分维度进行拆分;将数据量较大的制度文件拆分为数据量较小的中间文件,可以并行将中间文件拆分为多个制度子文件,进而可以提高文件拆分效率。
应该理解的是,虽然图2和图4的流程图中的各个步骤按照箭头的指示依次显示,但是这些步骤并不是必然按照箭头指示的顺序依次执行。除非本文中有明确的说明,这些步骤的执行并没有严格的顺序限制,这些步骤可以以其它 的顺序执行。而且,图2和图4中的至少一部分步骤可以包括多个子步骤或者多个阶段,这些子步骤或者阶段并不必然是在同一时刻执行完成,而是可以在不同的时刻执行,这些子步骤或者阶段的执行顺序也不必然是依次进行,而是可以与其它步骤或者其它步骤的子步骤或者阶段的至少一部分轮流或者交替地执行。
在其中一个实施例中,如图6所示,提供了一种制度信息处理装置,包括:信息扩展模块602、信息分类模块604和信息归档模块606,其中:
信息扩展模块602,用于监测终端发布的制度信息,对制度信息进行分词得到对应的原始词语集合;原始词语集合包括多个原始词语;对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;根据各个扩展词语集合形成制度信息对应的扩展制度信息集合。信息分类模块604,用于将扩展制度信息集合输入预设的制度管理模型,得到制度信息对应的目标类别。信息归档模块606,用于获取多个目标信息树分别对应的类别标注,筛选包含与目标类别对应类别标注的目标信息树,将制度信息添加至筛选得到的目标信息树。
在其中一个实施例中,制度信息包括制度描述信息;信息扩展模块602还用于检测制度描述信息是否包含类别信息;若包含,则根据类别信息将制度信息添加至相应的目标信息树;否则,对制度信息进行分词得到对应的原始词语集合。
在其中一个实施例中,该装置还包括模型训练模块608,用于获取训练样本数据;训练样本数据包括多个样本制度信息及分别对应的类别标注;对各个样本制度信息进行分词和同义扩展处理,得到每个样本制度信息分别对应的扩展样本制度信息集合;根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,得到制度管理模型。
在其中一个实施例中,扩展样本制度信息集合包括多组扩展样本制度信息;模型训练模块608还用于获取特征项,计算特征项在一组扩展样本制度信息的词频权重;计算特征项在整个训练样本数据中的文档频率;根据词频权重和文档频率计算特征项对应的特征权重;根据特征权重选择特征项作为相应扩展样本制度信息的特征词;根据特征词提取各个扩展样本标准信息的特征。
在其中一个实施例中,制度信息包括制度描述信息和制度文件;信息归档 模块606还用于根据制度描述信息生成信息节点;检测筛选得到的目标信息树中是否已存在相同的信息节点;若不存在,则将信息节点添加至相应的目标信息树,将制度文件关联至信息节点。
在其中一个实施例中,制度信息包括制度描述信息及关联的制度文件;制度文件包括多个制度条款以及分别对应的适用对象标识;关联信息树具有对应的适用对象标识;信息归档模块606还用于对制度文件进行拆分,利用每个适用对象标识对应的制度条款生成相应适用对象标识对应的制度子文件;获取目标信息树对应的多个关联信息树;根据适用对象标识,将制度描述信息及制度子文件添加至相应的关联信息树。
在其中一个实施例中,信息归档模块606还用于计算制度文件的数据量,检测数据量是否超过阈值;当数据量超过阈值时,获取预设的目标数据量,根据目标数据量确定制度文件的拆分位置;检测拆分位置是否位于相邻分隔符之间;当拆分位置位于一个分隔符处时,在拆分位置将制度文件拆分为多个中间文件;当拆分位置位于相邻分隔符之间时,在相邻分隔符中任意一个分隔符处将制度文件拆分为多个中间文件;按照预设的拆分规则,对多个中间文件进行拆分。
关于制度信息处理装置的具体限定可以参见上文中对于制度信息处理方法的限定,在此不再赘述。上述制度信息处理装置中的各个模块可全部或部分通过软件、硬件及其组合来实现。上述各模块可以硬件形式内嵌于或独立于计算机设备中的处理器中,也可以以软件形式存储于计算机设备中的存储器中,以便于处理器调用执行以上各个模块对应的操作。
在其中一个实施例中,提供了一种计算机设备,该计算机设备可以是服务器,其内部结构图可以如图7所示。该计算机设备包括通过系统总线连接的处理器、存储器、网络接口和数据库。其中,该计算机设备的处理器用于提供计算和控制能力。该计算机设备的存储器包括非易失性存储介质、内存储器。该非易失性存储介质存储有操作系统、计算机可读指令和数据库。该内存储器为非易失性存储介质中的操作系统和计算机可读指令的运行提供环境。该计算机设备的数据库用于存储制度信息。该计算机设备的网络接口用于与外部的终端通过网络连接通信。该计算机可读指令被处理器执行时实现本申请任意一个实施例中提供的制度信息处理方法的步骤。
本领域技术人员可以理解,图7中示出的结构,仅仅是与本申请方案相关 的部分结构的框图,并不构成对本申请方案所应用于其上的计算机设备的限定,具体的计算机设备可以包括比图中所示更多或更少的部件,或者组合某些部件,或者具有不同的部件布置。
一个或多个存储有计算机可读指令的非易失性存储介质,计算机可读指令被一个或多个处理器执行时,使得一个或多个处理器实现本申请任意一个实施例中提供的制度信息处理方法的步骤。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机可读指令来指令相关的硬件来完成,计算机可读指令可存储于一非易失性计算机可读取存储介质中,该计算机可读指令在执行时,可包括如上述各方法的实施例的流程。其中,本申请所提供的各实施例中所使用的对存储器、存储、数据库或其它介质的任何引用,均可包括非易失性和/或易失性存储器。非易失性存储器可包括只读存储器(ROM)、可编程ROM(PROM)、电可编程ROM(EPROM)、电可擦除可编程ROM(EEPROM)或闪存。易失性存储器可包括随机存取存储器(RAM)或者外部高速缓冲存储器。作为说明而非局限,RAM以多种形式可得,诸如静态RAM(SRAM)、动态RAM(DRAM)、同步DRAM(SDRAM)、双数据率SDRAM(DDRSDRAM)、增强型SDRAM(ESDRAM)、同步链路(Synchlink)DRAM(SLDRAM)、存储器总线(Rambus)直接RAM(RDRAM)、直接存储器总线动态RAM(DRDRAM)、以及存储器总线动态RAM(RDRAM)等。
以上实施例的各技术特征可以进行任意的组合,为使描述简洁,未对上述实施例中的各个技术特征所有可能的组合都进行描述,然而,只要这些技术特征的组合不存在矛盾,都应当认为是本说明书记载的范围。
以上实施例仅表达了本申请的几种实施方式,其描述较为具体和详细,但并不能因此而理解为对发明专利范围的限制。应当指出的是,对于本领域的普通技术人员来说,在不脱离本申请构思的前提下,还可以做出若干变形和改进,这些都属于本申请的保护范围。因此,本申请专利的保护范围应以所附权利要求为准。

Claims (20)

  1. 一种制度信息处理方法,包括:
    监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;
    对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;
    根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;
    将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;及
    获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
  2. 根据权利要求1所述的方法,其特征在于,所述制度信息包括制度描述信息;所述对所述制度信息进行分词得到对应的原始词语集合之前,所述方法还包括:
    检测所述制度描述信息是否包含类别信息;
    若包含,则根据所述类别信息将所述制度信息添加至相应的目标信息树;及
    否则,对所述制度信息进行分词得到对应的原始词语集合。
  3. 根据权利要求1所述的方法,其特征在于,所述制度管理模型的生成,包括:获取训练样本数据;所述训练样本数据包括多个样本制度信息及分别对应的类别标注;
    对各个所述样本制度信息进行分词和同义扩展处理,得到每个所述样本制度信息分别对应的扩展样本制度信息集合;及
    根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,得到所述制度管理模型。
  4. 根据权利要求3所述的方法,其特征在于,所述扩展样本制度信息集合包括多组扩展样本制度信息;根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,包括:
    获取特征项,计算所述特征项在一组所述扩展样本制度信息的词频权重;
    计算所述特征项在整个训练样本数据中的文档频率;
    根据所述词频权重和文档频率计算所述特征项对应的特征权重;
    根据所述特征权重选择所述特征项作为相应扩展样本制度信息的特征词;及
    根据所述特征词提取各个所述扩展样本标准信息的特征。
  5. 根据权利要求1所述的方法,其特征在于;所述制度信息包括制度描述信息和制度文件;所述将所述制度信息添加至筛选得到的目标信息树,包括:
    根据所述制度描述信息生成信息节点;
    检测筛选得到的目标信息树中是否已存在相同的信息节点;及
    若不存在,则将所述信息节点添加至相应的目标信息树,将所述制度文件关联至所述信息节点。
  6. 根据权利要求1所述的方法,其特征在于,所述制度信息包括制度描述信息及关联的制度文件;所述制度文件包括多个制度条款以及分别对应的适用对象标识;所述关联信息树具有对应的适用对象标识;所述方法还包括:
    对所述制度文件进行拆分,利用每个适用对象标识对应的制度条款生成相应适用对象标识对应的制度子文件;
    获取所述目标信息树对应的多个关联信息树;及
    根据所述适用对象标识,将所述制度描述信息及所述制度子文件添加至相应的关联信息树。
  7. 根据权利要求6所述的方法,其特征在于,所述对所述制度文件进行拆分,包括:
    计算所述制度文件的数据量,检测所述数据量是否超过阈值;
    当所述数据量超过阈值时,获取预设的目标数据量,根据所述目标数据量确定所述制度文件的拆分位置;
    检测所述拆分位置是否位于相邻分隔符之间;
    当所述拆分位置位于一个分隔符处时,在所述拆分位置将所述制度文件拆分为多个中间文件;
    当所述拆分位置位于相邻分隔符之间时,在所述相邻分隔符中任意一个分隔符处将所述制度文件拆分为多个中间文件;及
    按照预设的拆分规则,对多个所述中间文件进行拆分。
  8. 一种制度信息处理装置,包括:
    信息扩展模块,用于监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;
    信息分类模块,用于将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;
    信息归档模块,用于获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
  9. 一种计算机设备,包括存储器及一个或多个处理器,所述存储器中储存有计算机可读指令,所述计算机可读指令被所述一个或多个处理器执行时,使得所述一个或多个处理器执行以下步骤:
    监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;
    对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;
    根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;
    将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;及
    获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
  10. 根据权利要求9所述的计算机设备,其特征在于,所述制度信息包括制度描述信息;所述处理器执行所述计算机可读指令时还执行以下步骤:
    检测所述制度描述信息是否包含类别信息;
    若包含,则根据所述类别信息将所述制度信息添加至相应的目标信息树;及
    否则,对所述制度信息进行分词得到对应的原始词语集合。
  11. 根据权利要求9所述的计算机设备,其特征在于,所述处理器执行 所述计算机可读指令时还执行以下步骤:获取训练样本数据;所述训练样本数据包括多个样本制度信息及分别对应的类别标注;
    对各个所述样本制度信息进行分词和同义扩展处理,得到每个所述样本制度信息分别对应的扩展样本制度信息集合;及
    根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,得到所述制度管理模型。
  12. 根据权利要求11所述的计算机设备,其特征在于,所述扩展样本制度信息集合包括多组扩展样本制度信息;所述处理器执行所述计算机可读指令时还执行以下步骤:
    获取特征项,计算所述特征项在一组所述扩展样本制度信息的词频权重;
    计算所述特征项在整个训练样本数据中的文档频率;
    根据所述词频权重和文档频率计算所述特征项对应的特征权重;
    根据所述特征权重选择所述特征项作为相应扩展样本制度信息的特征词;及
    根据所述特征词提取各个所述扩展样本标准信息的特征。
  13. 根据权利要求9所述的计算机设备,其特征在于,所述制度信息包括制度描述信息和制度文件;所述处理器执行所述计算机可读指令时还执行以下步骤:
    根据所述制度描述信息生成信息节点;
    检测筛选得到的目标信息树中是否已存在相同的信息节点;及
    若不存在,则将所述信息节点添加至相应的目标信息树,将所述制度文件关联至所述信息节点。
  14. 根据权利要求9所述的计算机设备,其特征在于,所述制度信息包括制度描述信息及关联的制度文件;所述制度文件包括多个制度条款以及分别对应的适用对象标识;所述关联信息树具有对应的适用对象标识;所述处理器执行所述计算机可读指令时还执行以下步骤:
    对所述制度文件进行拆分,利用每个适用对象标识对应的制度条款生成相应适用对象标识对应的制度子文件;
    获取所述目标信息树对应的多个关联信息树;及
    根据所述适用对象标识,将所述制度描述信息及所述制度子文件添加至相应的关联信息树。
  15. 一个或多个存储有计算机可读指令的非易失性计算机可读存储介质,所述计算机可读指令被一个或多个处理器执行时,使得所述一个或多个处理器执行以下步骤:
    监测终端发布的制度信息,对所述制度信息进行分词得到对应的原始词语集合;所述原始词语集合包括多个原始词语;
    对各个原始词语进行同义扩展,生成每个原始词语对应的扩展词语集合;
    根据各个扩展词语集合形成所述制度信息对应的扩展制度信息集合;
    将所述扩展制度信息集合输入预设的制度管理模型,得到所述制度信息对应的目标类别;及
    获取多个目标信息树分别对应的类别标注,筛选包含与所述目标类别对应类别标注的目标信息树,将所述制度信息添加至筛选得到的目标信息树。
  16. 根据权利要求15所述的存储介质,其特征在于,所述制度信息包括制度描述信息;所述计算机可读指令被所述处理器执行时还执行以下步骤:
    检测所述制度描述信息是否包含类别信息;
    若包含,则根据所述类别信息将所述制度信息添加至相应的目标信息树;及
    否则,对所述制度信息进行分词得到对应的原始词语集合。
  17. 根据权利要求15所述的存储介质,其特征在于,所述计算机可读指令被所述处理器执行时还执行以下步骤:获取训练样本数据;所述训练样本数据包括多个样本制度信息及分别对应的类别标注;
    对各个所述样本制度信息进行分词和同义扩展处理,得到每个所述样本制度信息分别对应的扩展样本制度信息集合;及
    根据各个扩展样本制度信息集合和对应的类别标注,通过支持向量机算法对初始的制度管理模型进行训练,得到所述制度管理模型。
  18. 根据权利要求17所述的存储介质,其特征在于,所述扩展样本制度信息集合包括多组扩展样本制度信息;所述计算机可读指令被所述处理器执行时还执行以下步骤:
    获取特征项,计算所述特征项在一组所述扩展样本制度信息的词频权重;
    计算所述特征项在整个训练样本数据中的文档频率;
    根据所述词频权重和文档频率计算所述特征项对应的特征权重;
    根据所述特征权重选择所述特征项作为相应扩展样本制度信息的特征词;及
    根据所述特征词提取各个所述扩展样本标准信息的特征。
  19. 根据权利要求15所述的存储介质,其特征在于,所述制度信息包括制度描述信息和制度文件;所述计算机可读指令被所述处理器执行时还执行以下步骤:
    根据所述制度描述信息生成信息节点;
    检测筛选得到的目标信息树中是否已存在相同的信息节点;及
    若不存在,则将所述信息节点添加至相应的目标信息树,将所述制度文件关联至所述信息节点。
  20. 根据权利要求15所述的存储介质,其特征在于,所述制度信息包括制度描述信息及关联的制度文件;所述制度文件包括多个制度条款以及分别对应的适用对象标识;所述关联信息树具有对应的适用对象标识;所述计算机可读指令被所述处理器执行时还执行以下步骤:
    对所述制度文件进行拆分,利用每个适用对象标识对应的制度条款生成相应适用对象标识对应的制度子文件;
    获取所述目标信息树对应的多个关联信息树;及
    根据所述适用对象标识,将所述制度描述信息及所述制度子文件添加至相应的关联信息树。
PCT/CN2018/095549 2018-04-09 2018-07-13 制度信息处理方法、装置、计算机设备和存储介质 Ceased WO2019196228A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN201810313040.X 2018-04-09
CN201810313040.XA CN108509424B (zh) 2018-04-09 2018-04-09 制度信息处理方法、装置、计算机设备和存储介质

Publications (1)

Publication Number Publication Date
WO2019196228A1 true WO2019196228A1 (zh) 2019-10-17

Family

ID=63380978

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2018/095549 Ceased WO2019196228A1 (zh) 2018-04-09 2018-07-13 制度信息处理方法、装置、计算机设备和存储介质

Country Status (2)

Country Link
CN (1) CN108509424B (zh)
WO (1) WO2019196228A1 (zh)

Cited By (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111858923A (zh) * 2019-12-24 2020-10-30 北京嘀嘀无限科技发展有限公司 一种文本分类方法、系统、装置及存储介质
CN112199951A (zh) * 2020-11-04 2021-01-08 支付宝(杭州)信息技术有限公司 一种事件信息生成的方法及装置
CN112307174A (zh) * 2020-11-20 2021-02-02 深圳壹账通创配科技有限公司 多平台数据整合方法、装置、计算机设备及可读存储介质
CN113033665A (zh) * 2021-03-26 2021-06-25 北京沃东天骏信息技术有限公司 样本扩展方法、训练方法和系统、及样本学习系统
CN113254398A (zh) * 2020-12-29 2021-08-13 深圳市怡化时代科技有限公司 样本文件管理方法、装置、设备和介质
CN117689354A (zh) * 2024-02-04 2024-03-12 芯知科技(江苏)有限公司 基于云服务的招聘信息的智能处理方法及平台

Families Citing this family (6)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN109582788A (zh) * 2018-11-09 2019-04-05 北京京东金融科技控股有限公司 垃圾评论训练、识别方法、装置、设备及可读存储介质
CN110162607B (zh) * 2019-02-20 2021-08-31 北京捷风数据技术有限公司 一种基于卷积神经网络的政府组织公文信息追溯方法及装置
CN111198945A (zh) * 2019-12-03 2020-05-26 泰康保险集团股份有限公司 数据处理方法、装置、介质及电子设备
CN111428967A (zh) * 2020-03-02 2020-07-17 四川宝石花鑫盛油气运营服务有限公司 基于岗位为基本单元的文件管理方法及装置
CN112417887B (zh) * 2020-11-20 2023-12-05 小沃科技有限公司 敏感词句识别模型处理方法、及其相关设备
CN121578930A (zh) * 2025-10-29 2026-02-27 北京百融睿博科技有限公司 一种文件的读取方法及装置

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104978328A (zh) * 2014-04-03 2015-10-14 北京奇虎科技有限公司 一种获取层级分类器以及文本分类的方法及装置
CN107436875A (zh) * 2016-05-25 2017-12-05 华为技术有限公司 文本分类方法及装置
CN107862046A (zh) * 2017-11-07 2018-03-30 宁波爱信诺航天信息有限公司 一种基于短文本相似度的税务商品编码分类方法及系统

Family Cites Families (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US4611298A (en) * 1983-06-03 1986-09-09 Harding And Harris Behavioral Research, Inc. Information storage and retrieval system and method
CN100538695C (zh) * 2004-07-22 2009-09-09 国际商业机器公司 构造、维护个性化分类树的方法及系统
JP5112027B2 (ja) * 2007-11-29 2013-01-09 株式会社日立ソリューションズ 文書群提示装置および文書群提示プログラム
CN106991112B (zh) * 2016-11-07 2021-05-25 创新先进技术有限公司 信息查询方法及装置
CN107463715A (zh) * 2017-09-13 2017-12-12 电子科技大学 基于信息增益的英文社交媒体账号分类方法

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN104978328A (zh) * 2014-04-03 2015-10-14 北京奇虎科技有限公司 一种获取层级分类器以及文本分类的方法及装置
CN107436875A (zh) * 2016-05-25 2017-12-05 华为技术有限公司 文本分类方法及装置
CN107862046A (zh) * 2017-11-07 2018-03-30 宁波爱信诺航天信息有限公司 一种基于短文本相似度的税务商品编码分类方法及系统

Cited By (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN111858923A (zh) * 2019-12-24 2020-10-30 北京嘀嘀无限科技发展有限公司 一种文本分类方法、系统、装置及存储介质
CN112199951A (zh) * 2020-11-04 2021-01-08 支付宝(杭州)信息技术有限公司 一种事件信息生成的方法及装置
CN112307174A (zh) * 2020-11-20 2021-02-02 深圳壹账通创配科技有限公司 多平台数据整合方法、装置、计算机设备及可读存储介质
CN113254398A (zh) * 2020-12-29 2021-08-13 深圳市怡化时代科技有限公司 样本文件管理方法、装置、设备和介质
CN113033665A (zh) * 2021-03-26 2021-06-25 北京沃东天骏信息技术有限公司 样本扩展方法、训练方法和系统、及样本学习系统
CN117689354A (zh) * 2024-02-04 2024-03-12 芯知科技(江苏)有限公司 基于云服务的招聘信息的智能处理方法及平台
CN117689354B (zh) * 2024-02-04 2024-04-19 芯知科技(江苏)有限公司 基于云服务的招聘信息的智能处理方法及平台

Also Published As

Publication number Publication date
CN108509424A (zh) 2018-09-07
CN108509424B (zh) 2021-08-10

Similar Documents

Publication Publication Date Title
WO2019196228A1 (zh) 制度信息处理方法、装置、计算机设备和存储介质
US9792289B2 (en) Systems and methods for file clustering, multi-drive forensic analysis and data protection
WO2019196226A1 (zh) 制度信息查询方法、装置、计算机设备和存储介质
US20250165521A1 (en) Automatic Detection and Transfer of Relevant Image Data to Content Collections
Shahana et al. Evaluation of features on sentimental analysis
WO2020207167A1 (zh) 文本分类方法、装置、设备及计算机可读存储介质
WO2019196224A1 (zh) 制度信息处理方法、装置、计算机设备和存储介质
JP2025508358A (ja) セキュリティインシデントを検出するために変則的なコンピュータイベントを識別するための方法およびシステム
CN111639178A (zh) 生命科学文档的自动分类和解释
WO2019196219A1 (zh) 制度信息安全监控方法、装置、计算机设备和存储介质
US11176403B1 (en) Filtering detected objects from an object recognition index according to extracted features
CN114724156B (zh) 表单识别方法、装置及电子设备
WO2022105119A1 (zh) 意图识别模型的训练语料生成方法及其相关设备
JP2019057068A (ja) 情報処理装置およびコンピュータプログラム
CN119903023A (zh) 基于大模型的文件元数据的生成方法、装置、设备和介质
CN112214615A (zh) 基于知识图谱的政策文件处理方法、装置和存储介质
Sarmento et al. Automatic extraction of quotes and topics from news feeds
WO2025066156A1 (zh) 一种解释多组黑盒人工智能模型之间的公共交互效用的方法和系统
KR101019627B1 (ko) 패턴 기반 참고문헌 자동 구축 시스템 및 방법과 이를 위한기록매체
Ła̧giewka et al. Distributed image retrieval with colour and keypoint features
Ruba et al. Building a custom sentiment analysis tool based on an ontology for Twitter posts
Yin et al. Content‐Based Image Retrial Based on Hadoop
Thepade et al. Decision fusion-based approach for content-based image classification
CN119067103A (zh) 电力应急预案文本生成方法、装置、计算机设备和存储介质
US12316678B2 (en) Security audit of data-at-rest

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 18914005

Country of ref document: EP

Kind code of ref document: A1

NENP Non-entry into the national phase

Ref country code: DE

32PN Ep: public notification in the ep bulletin as address of the adressee cannot be established

Free format text: NOTING OF LOSS OF RIGHTS PURSUANT TO RULE 112(1) EPC (EPO FORM 1205A DATED 26.01.2021)

122 Ep: pct application non-entry in european phase

Ref document number: 18914005

Country of ref document: EP

Kind code of ref document: A1