WO2025201470A1 - 模型训练数据生成 - Google Patents

模型训练数据生成

Info

Publication number
WO2025201470A1
WO2025201470A1 PCT/CN2025/085413 CN2025085413W WO2025201470A1 WO 2025201470 A1 WO2025201470 A1 WO 2025201470A1 CN 2025085413 W CN2025085413 W CN 2025085413W WO 2025201470 A1 WO2025201470 A1 WO 2025201470A1
Authority
WO
WIPO (PCT)
Prior art keywords
domain data
domain
feature
model
data
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
PCT/CN2025/085413
Other languages
English (en)
French (fr)
Inventor
刘超
吴俊�
朱磊
陈国利
孙志铭
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Alipay Hangzhou Information Technology Co Ltd
Original Assignee
Alipay Hangzhou Information Technology Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Alipay Hangzhou Information Technology Co Ltd filed Critical Alipay Hangzhou Information Technology Co Ltd
Publication of WO2025201470A1 publication Critical patent/WO2025201470A1/zh
Pending legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/22Matching criteria, e.g. proximity measures
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/25Fusion techniques

Definitions

  • This specification provides a method, device, storage medium and electronic device for generating model training data, and the technical solution is as follows.
  • this specification provides an electronic device that may include a processor and a memory; wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned method steps.
  • this specification provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by a processor to execute any one of the above method steps.
  • the beneficial effects brought about by the technical solutions provided in some embodiments of this specification include at least: by determining the domain data features corresponding to the target domain data and the model processing labels of the domain data features, the domain data features with large differences from other domain data features are determined based on the domain data features, and the model processing labels corresponding to the domain data features with large differences from other domain data features are often prone to labeling errors.
  • the model processing labels of the domain data features that are prone to labeling errors in the domain data features are updated, thereby obtaining a target domain model training data set, improving the training sample quality of the target domain data feature set, and finally using the obtained target domain model training data set to train the generative processing model, thereby obtaining a high-quality processing model for the target domain.
  • FIG1 is a schematic diagram of a scenario of a model training data generation system provided in this specification.
  • FIG2 is a flow chart of a method for generating model training data provided in an embodiment of this specification.
  • FIG3 is a schematic diagram of a process for determining a target domain model training dataset according to an embodiment of this specification.
  • FIG4 is a schematic diagram of a process for determining domain-specific data features according to an embodiment of this specification.
  • FIG5 is a schematic diagram of another process for determining domain-specific data features according to an embodiment of this specification.
  • FIG6 is a schematic diagram of another process for determining domain-specific data features according to an embodiment of this specification.
  • FIG7 is a schematic diagram of a process for determining a new specific domain data feature in a domain data feature set according to an embodiment of this specification.
  • FIG8 is a schematic diagram of a process for determining domain data features corresponding to target domain data according to an embodiment of this specification.
  • FIG9 is a flow chart of a generative processing model and target domain data provided by an embodiment of this specification.
  • FIG10 is a schematic diagram of a process for updating a target domain model training dataset provided in an embodiment of this specification.
  • FIG11 is a model training data generating device provided in an embodiment of this specification.
  • FIG12 is a structural block diagram of an electronic device provided in an embodiment of this specification.
  • FIG13 is a schematic diagram of the structure of an operating system and user space provided in an embodiment of this specification.
  • FIG14 is an architecture diagram of the Android operating system in FIG13 provided in an embodiment of this specification.
  • FIG15 is an architectural diagram of the IOS operating system in FIG13 provided in an embodiment of this specification.
  • Figure 1 is a schematic diagram of a scenario of a model training data generation system provided in this specification.
  • the model training data generation system may include at least a client cluster and a service platform 100.
  • the client cluster may include at least one client, as shown in FIG1 , specifically including client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
  • Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem.
  • Electronic devices may be referred to by different names in different networks, such as user equipment, access terminal, subscriber unit, subscriber station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication device, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolution networks.
  • PDA personal digital assistant
  • the service platform 100 can be a separate server device, such as a rack-mounted, blade, tower, or cabinet-mounted server device, or a workstation, mainframe computer, or other hardware device with strong computing capabilities; it can also be a server cluster composed of multiple servers.
  • the servers in the service cluster can be symmetrically composed, wherein each server has equivalent functions and status in the transaction link, and each server can provide services to the outside world independently.
  • the independent service can be understood as not requiring the assistance of other servers.
  • the service platform 100 may establish a communication connection with at least one client in the client cluster, and complete data interaction during the model training data generation process based on the communication connection, such as online transaction data interaction.
  • the service platform 100 may implement model training for the generative processing model based on the target domain model training data set obtained based on the model training data generation method of the present specification.
  • the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication.
  • the network can be a wireless network or a wired network.
  • Wireless networks include, but are not limited to, cellular networks, wireless local area networks, infrared networks, or Bluetooth networks.
  • Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area network.
  • technologies and/or formats including Hypertext Markup Language (HTML) and Extensible Markup Language (XML) are used to represent data (such as target compressed packages) exchanged over the network.
  • SSL Secure Sockets Layer
  • TLS Transport Layer Security
  • VPN Virtual Private Network
  • IPsec Internet Protocol Security
  • SSL Secure Sockets Layer
  • TLS Transport Layer Security
  • VPN Virtual Private Network
  • IPsec Internet Protocol Security
  • customized and/or dedicated data communication technologies can be used to replace or supplement the above-mentioned data communication technologies.
  • the model training data generation system embodiments provided in this specification share the same concept as the model training data generation method described in one or more embodiments.
  • the execution entity corresponding to the model training data generation method described in one or more embodiments of the specification may be the aforementioned service platform 100; the execution entity corresponding to the model training data generation method described in one or more embodiments of the specification may also be the electronic device corresponding to the client, specifically determined based on the actual application environment.
  • the implementation process of the model training data generation system embodiment can be found in the following method embodiment, which will not be detailed here.
  • model training data generation method provided by one or more embodiments of this specification is introduced in detail below.
  • Figure 2 is a flow chart illustrating a method for generating model training data according to an embodiment of this specification.
  • This method can be implemented using a computer program and can be run on a model training data generation device based on the von Neumann architecture.
  • the computer program can be integrated into an application or run as a standalone tool application.
  • the model training data generation device can be a service platform.
  • the model training data generation method includes the following steps.
  • the generative processing model can be a generative large language model, which is mainly constructed using deep learning technologies such as neural networks. This type of model can learn based on large amounts of text data and combine complex algorithms to understand, interpret and generate human language.
  • generative processing models can be applied to target domains to process data.
  • generative processing models are effective at general language processing tasks, but their performance in specialized verticals is limited, as these domains are often filled with complex terminology and unique operational rules.
  • the target domain refers to a specific problem domain or knowledge domain. It can be a collection of domain-specific terminology, concepts, rules, processes, and entities.
  • the target domain can be any specific knowledge domain, such as finance, healthcare, education, or e-commerce, as well as any specific problem domain, specifically an application domain, such as natural language processing, image recognition, or machine learning.
  • the target domain data can be of various types, such as structured data, unstructured data, semi-structured data, etc.
  • Structured data refers to data with a clear format and rules, such as tabular data and relational data in a database; unstructured data refers to data without a clear format and rules, such as text data, image data, audio data, etc.
  • Semi-structured data is data between structured data and unstructured data, such as XML (Extensible Markup Language) documents and JSON (JavaScript Object Notation) data.
  • Target domain data can also be from finance, healthcare, education, or transportation.
  • financial data in the financial sector can include securities market data, exchange rate data, financial institution data, and financial product data.
  • Medical data can include patient data, drug data, medical device data, and disease diagnosis data. Data from different domains has its own unique characteristics and application scenarios, so the specificity of each domain must be considered during data processing and analysis.
  • S204 Determine domain data features corresponding to the target domain data and model processing labels of the domain data features.
  • domain data features are extracted from the target domain data to obtain domain data features.
  • the domain data features include both general domain dimension features and domain vertical dimension features.
  • domain data features include but are not limited to domain data feature vectors, domain data feature matrices, domain data feature graphs, domain data feature values, etc. The selection of these presentation forms depends on the specific application scenario.
  • a domain data feature vector is a concise representation of domain data features, representing data attributes or characteristics. For example, in text classification tasks, a domain data feature vector can represent keywords, phrases, or themes in a text.
  • a domain data feature matrix is a more detailed representation of domain data features.
  • a domain data feature graph is an intuitive representation of domain data features, often used to visualize the relationships and network structure between domain data features.
  • data features of two dimensions can be extracted from the target domain data, namely, the general domain dimension of the target domain data and the domain vertical dimension corresponding to the target domain.
  • the data features of these two dimensions are then fused to obtain domain data features.
  • domain data features are domain data feature vectors
  • data feature extraction includes a data vectorization process
  • domain data features are domain data feature matrices
  • data feature extraction includes a data matrixization process
  • domain data features are domain data feature graphs
  • data feature extraction includes a data visualization process
  • data feature extraction includes a numeration process.
  • the generative processing model can process general processing tasks more efficiently and accurately after learning the characteristics of the general domain dimension; and, after learning the characteristics of the domain vertical dimension, the generative processing model can process the processing tasks of the target domain more efficiently and accurately based on the characteristics of the general domain dimension.
  • domain data features that differ significantly from other domain data features are identified based on feature similarity between the domain data features. Since model processing labels corresponding to domain data features that differ significantly from other domain data features are often prone to mislabeling, model processing labels for domain data features that are prone to mislabeling can be updated based on the domain data features.
  • the specific model processing labels corresponding to the specific domain data features can be modified to obtain modified model processing labels.
  • the label modification process can be based on the training model; or, the label modification process can be performed by expert intervention.
  • all feature similarities of each field data feature set are obtained, that is, the feature similarities between each two field data features in each field data feature set, and then statistical analysis and processing are performed based on all the obtained feature similarities to obtain feature set similarity distribution information.
  • the feature set similarity distribution information may be domain data feature pairs distributed in each feature similarity interval.
  • domain data feature pairs distributed in a feature similarity interval of 40% to 50% and their corresponding feature similarities and domain data feature pairs distributed in a feature similarity interval of 70% to 80% and their corresponding feature similarities.
  • S406 Determine specific domain data features from the domain data feature set based on the feature set similarity distribution information.
  • the feature set similarity distribution information records the frequency of occurrence of data features in each field in each feature similarity interval.
  • domain data features that appear more frequently in the low feature similarity area can be considered as vectors with lower similarity to other domain data features, and therefore can be used as specific domain data features.
  • clustering is performed by the feature similarity between two domain data features to obtain at least two types of domain data feature sets, and then the feature set similarity distribution information in each domain data feature set is determined, so as to quickly determine specific domain data features with low similarity to other domain data features through the feature set similarity distribution information.
  • determining specific domain data features from the domain data feature set based on feature set similarity distribution information includes the following steps.
  • S502 Based on the feature set similarity distribution information, query the target feature similarity whose feature similarity is less than or equal to the feature similarity threshold from the domain data feature set, determine the domain data feature pairs corresponding to the target feature similarity, and determine the domain data feature pair set based on the domain data feature pairs.
  • the feature similarity threshold can be flexibly set based on the feature set similarity distribution information. For example, if the minimum feature similarity interval in the feature set similarity distribution information is 30% to 40%, the feature similarity threshold can be set to 40% or 50%. Of course, the feature similarity threshold can also be manually pre-set. The feature similarity thresholds corresponding to different domain data feature sets can be the same or different.
  • the target feature similarities whose feature similarity is less than or equal to the feature similarity threshold in the feature set similarity distribution information are searched from the domain data feature set.
  • Each target feature similarity corresponds to a pair of domain data features, i.e., a domain data feature pair.
  • the domain data feature pairs corresponding to each target feature similarity are put into the same set to obtain the domain data feature pair set corresponding to the domain data feature set.
  • S504 Determine the feature repetition frequency of the domain data feature for each reference domain data feature in the set.
  • the feature similarity of each domain data feature pair in the domain data feature pair set is less than or equal to the feature similarity threshold. Therefore, the vector similarity of each domain data feature pair in the domain data feature pair set is low. At this time, the feature repetition frequency of each reference domain data feature in the domain data feature pair set can be counted.
  • the higher the feature repetition frequency of the reference domain data feature the greater the difference between the reference domain data feature and other reference domain data features; the lower the feature repetition frequency of the reference domain data feature, the smaller the difference between the reference domain data feature and other reference domain data features.
  • S506 Determine the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
  • the reference domain data feature when the feature repetition frequency of the reference domain data feature is higher, the reference domain data feature is more different from other reference domain data features.
  • the reference domain data feature can be considered as a specific domain data feature in the domain data feature set.
  • the specific domain data feature can be one or more.
  • the feature repetition frequencies of each reference domain data feature may be sorted, such as in descending order, and the reference domain data feature corresponding to the feature repetition frequency ranked first may be used as the specific domain data feature in the domain data feature set.
  • domain data feature pairs with higher similarity in the domain data feature set are filtered out by using a feature similarity threshold to obtain domain data feature pairs with lower similarity in the domain data feature set, i.e., a set of domain data feature pairs. Then, the feature repetition frequency of each reference domain data feature in the set of domain data feature pairs is counted. When the feature repetition frequency of the reference domain data feature is higher, it indicates that the reference domain data feature is more different from other reference domain data features. At this time, the reference domain data feature can be considered as a specific domain data feature in the domain data feature set.
  • S602 Obtain an indicator ratio mapping relationship between a model training indicator and a label correction ratio, determine a target model training indicator for the generative processing model, and determine a target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship.
  • the model training metric is the expected performance metric after generative processing model training.
  • Different label correction ratios correspond to different model training metrics.
  • a higher label correction ratio results in a higher performance metric after generative processing model training;
  • a lower label correction ratio results in a lower performance metric after generative processing model training.
  • the label correction ratio is the proportion of specific domain data features in the domain data feature set that undergo label correction.
  • the target label correction ratio can be expressed as Stratify(D;P), where P is the target model training metric and Stratify can be a metric ratio mapping function.
  • the target model training indicator of the generative processing model is obtained. For example, when processing the target domain task in the target domain, the accuracy is 95%.
  • the indicator ratio mapping relationship can be used to determine the target label correction ratio corresponding to the target model training indicator.
  • S604 Determine the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set.
  • the target label correction ratio is multiplied by the number of domain data features in the domain data feature set to calculate the number of features that need to be label corrected, that is, the number of specific domain data features.
  • S606 Determine the specific domain data features in the domain data feature set based on the number of specific domain data features and the feature repetition frequency of each reference domain data feature.
  • step S404 the feature repetition frequencies of each reference domain data feature have been obtained, and these feature repetition frequencies can be sorted in descending order.
  • the number of specific domain data features is X
  • the reference domain data features corresponding to the first X feature repetition frequencies in the descending order can be used as the specific domain data features, thereby obtaining the specific domain data features in the domain data feature set.
  • the specific domain data features can be represented as the intersection of Cluster(D;K) and Stratify(D;P).
  • the embodiments of this specification determine the target label correction ratio through the mapping relationship between the target model training indicators and the indicator ratio, thereby determining the number of specific domain data features, and then screening the specific domain data features based on the feature repetition frequency of each reference domain data feature and the number of specific domain data features, thereby ensuring the training quality of the generative processing model while reducing the workload of label correction processing and improving the efficiency of model training data generation.
  • a target domain model training dataset is used to perform model training on a generative processing model, including: performing a first model training process on the generative processing model based on the target domain model training dataset to obtain a first generative processing model; detecting whether the first generative processing model meets a model training end condition; if the first generative processing model does not meet the model training end condition, determining to add a new specific domain data feature to the domain data feature set, performing label correction on the specific model processing label carried by the new specific domain data feature to obtain a new specific domain data feature after label correction, updating the target domain model training dataset based on the new specific domain data feature to obtain an updated target domain model training dataset, performing a second model training process on the first generative processing model based on the target domain model training dataset to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model training end condition; if the first generative processing model meets
  • the target domain model training dataset can be used to perform a first model training process on the generative processing model to obtain a first generative processing model.
  • the first generative processing model may not meet the model end training conditions, so the vector portion of the target domain model training dataset excluding the specific domain data features can be label-corrected again to train the first generative processing model again.
  • This process is then repeated, that is, incrementally training the model obtained in the previous round using each round of labeled data until the obtained model meets the model end training conditions. This determines that the model training is terminated and the target generative processing model is obtained.
  • Mi +1 Train( Mi , Si ).
  • Mi +1 is the model after the i+1th round of iteration, i is an integer greater than or equal to 0; Mi is the model of the i-th round; Si is the target domain model training dataset after the i-th round update; Train is the ongoing training process.
  • this solution not only significantly reduces the annotation burden of label correction processing, but also improves the accuracy of domain data annotation through multiple rounds of model training.
  • This optimizes the quality and availability of target domain model training datasets, laying a solid foundation for domain-specific model training. This enables the trained target generative processing model to accurately and efficiently handle target domain tasks in the target domain.
  • the model is continuously refined using newly identified annotated data, gradually reducing the need for high-quality annotations and optimizing the entire annotation process.
  • FIG. 7 is a schematic diagram of a process for determining a new specific domain data feature in a domain data feature set according to an embodiment of this specification.
  • determining a new specific domain data feature in a domain data feature set includes the following steps.
  • S702 Determine the current model training index of the first generative processing model, and determine the number of newly added specific domain data features based on the index difference between the current model training index and the target model training index.
  • the data labeling stage includes automated labeling processing, label correction processing, and cyclic label correction processing for newly added specific domain data features.
  • the target domain model training data set can be stored.
  • the domain knowledge base stage includes calculating data similarity, calculating data probability distribution information, the target domain model training data set and the corresponding model processing labels.
  • Figure 11 shows a model training data generation device provided in an embodiment of this specification.
  • the model training data generation device provided in this specification is described in detail. It should be noted that the model training data generation device shown in Figure 11 is used to execute the method of the embodiment shown in Figures 1 to 10 of this specification. For ease of explanation, only the parts related to this specification are shown. For specific technical details not disclosed, please refer to the embodiment shown in Figures 1 to 10 of this specification.
  • the update module 13 includes: a correction unit, suitable for determining specific domain data features based on domain data features, performing label correction processing on specific model processing labels corresponding to specific domain data features, and obtaining corrected model processing labels; an update unit, suitable for performing label update processing on model processing labels of domain data features based on corrected model processing labels, and obtaining a target domain model training data set.
  • the correction unit is further adapted to perform clustering processing on the domain data features to obtain at least two types of domain data feature sets, and determine the specific domain data features based on the domain data feature sets.
  • the correction unit includes: a calculation subunit, suitable for calculating the feature similarity between two domain data features, clustering based on the feature similarity, and obtaining at least two categories of domain data feature sets; a first determination subunit, suitable for determining the feature set similarity distribution information based on all feature similarities of the domain data feature set; a second determination subunit, suitable for determining specific domain data features from the domain data feature set based on the feature set similarity distribution information, and determining specific domain data features based on the specific domain data features.
  • the second determination subunit includes: a domain data feature pair set determination subunit, adapted to query target feature similarities whose feature similarities are less than or equal to a feature similarity threshold from the domain data feature set based on feature set similarity distribution information, determine the domain data feature pairs corresponding to the target feature similarities, and determine the domain data feature pair set based on the domain data feature pairs; a feature repetition frequency determination subunit, adapted to determine the feature repetition frequency of each reference domain data feature in the domain data feature pair set; and a domain data feature determination subunit, adapted to determine specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
  • the domain data feature determination subunit includes: an acquisition subunit, suitable for obtaining the indicator ratio mapping relationship between the model training indicator and the label correction ratio, determining the target model training indicator for the generative processing model, and determining the target label correction ratio based on the target model training indicator and the indicator ratio mapping relationship; a specific domain data feature quantity determination subunit, suitable for determining the number of specific domain data features based on the target label correction ratio and the number of domain data features in the domain data feature set; a specific domain data feature determination subunit, suitable for determining the specific domain data features in the domain data feature set based on the number of specific domain data features and the feature repetition frequency of each reference domain data feature.
  • the training module 14 includes: a first model training unit, suitable for performing a first model training process on the generative processing model based on the target domain model training data set to obtain a first generative processing model; a detection unit, suitable for detecting whether the first generative processing model meets the model end training condition; a first judgment unit, suitable for determining to add new specific domain data features to the domain data feature set if the first generative processing model does not meet the model end training condition, performing label correction on the specific model processing label carried by the new specific domain data feature to obtain the new specific domain data feature after label correction, updating the target domain model training data set based on the new specific domain data feature to obtain an updated target domain model training data set, performing a second model training process on the first generative processing model based on the target domain model training data set to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model end training condition; a second judgment unit, suitable for determining to end
  • the first judgment unit includes: a subunit for determining the number of newly added specific domain data features, suitable for determining the current model training index of the first generative processing model, and determining the number of newly added specific domain data features based on the index difference between the current model training index and the target model training index; a subunit for determining newly added specific domain data features, suitable for determining the newly added specific domain data features in the domain data feature set based on the number of newly added specific domain data features and the feature repetition frequency of each reference domain data feature.
  • the determination module 12 includes: an extraction unit, suitable for extracting the domain comprehensive features and domain attribute features of the target domain data, and extracting the general domain features of the target domain data; a domain data feature determination unit, suitable for determining the domain vertical features based on the domain comprehensive features and domain attribute features, and determining the domain data features based on the domain data features and the general domain data features.
  • F L is the domain vertical feature
  • is the first weight coefficient
  • is the second weight coefficient
  • f c is the domain comprehensive feature
  • f b is the domain attribute feature
  • the weight coefficient determination subunit is suitable for determining the third weight coefficient corresponding to the general domain data feature, and the sum of the first weight coefficient, the second weight coefficient and the third weight coefficient is 1
  • the domain data feature subunit is suitable for using the second calculation formula to obtain the domain data feature by using the domain vertical feature, the general domain data feature and the third weight coefficient, and determining the domain data feature based on the domain data feature
  • F D is the domain data feature
  • fi is the general domain data feature
  • is the third weight coefficient
  • the acquisition module 11 includes: a generative processing model acquisition unit, suitable for acquiring a generative processing model for the target domain; an identification unit, suitable for acquiring source data and a target domain identifier for the target domain, and identifying the target domain source data from the source data based on the target domain identifier; a target domain entity information determination unit, suitable for acquiring target domain entity definition information, and extracting entity information from the target domain source data based on the target domain entity definition information to obtain target domain entity information; a target domain data extraction unit, suitable for acquiring data extraction logic corresponding to the target domain entity information, and extracting data from the target domain entity information based on the data extraction logic to obtain target domain data.
  • the model training data generating device 1 also includes: a feature similarity determination module, suitable for determining the feature similarity between pairwise domain data features in the target domain model training data set; a reference feature pair determination module, suitable for determining a reference feature similarity whose feature similarity is greater than or equal to a similarity threshold and a reference feature pair corresponding to the reference feature similarity; a data fusion module, suitable for determining two similar domain data features corresponding to the reference feature pair, and performing data fusion processing on the similar domain data features to update the target domain model training data set.
  • a feature similarity determination module suitable for determining the feature similarity between pairwise domain data features in the target domain model training data set
  • a reference feature pair determination module suitable for determining a reference feature similarity whose feature similarity is greater than or equal to a similarity threshold and a reference feature pair corresponding to the reference feature similarity
  • a data fusion module suitable for determining two similar domain data features corresponding to the reference feature pair, and performing data fusion processing on the similar domain data features
  • model training data generation device only uses the division of the above functional modules as an example when executing the model training data generation method.
  • the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
  • the model training data generation device provided in the above embodiment and the model training data generation method embodiment are of the same concept. The implementation process is detailed in the method embodiment and will not be repeated here.
  • This specification also provides a computer storage medium, which can store multiple instructions.
  • the instructions are suitable for being loaded by a processor and executed by the model training data generation method of the embodiments shown in Figures 1 to 10 above.
  • the specific execution process can be found in the specific description of the embodiments shown in Figures 1 to 10, which will not be repeated here.
  • This specification also provides a computer program product, which stores at least one instruction, and the at least one instruction is loaded by the processor and executed by the model training data generation method of the embodiments shown in Figures 1 to 10 above.
  • the specific execution process can be found in the specific description of the embodiments shown in Figures 1 to 10, and will not be repeated here.
  • FIG 12 is a block diagram of the structure of an electronic device provided in an embodiment of this specification.
  • the electronic device described in this specification may include one or more of the following components: a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150.
  • the processor 110, the memory 120, the input device 130, and the output device 140 may be connected via the bus 150.
  • the processor 110 may include one or more processing cores.
  • the processor 110 utilizes various interfaces and circuits to connect various components within the electronic device. It executes instructions, programs, code sets, or instruction sets stored in the memory 120, as well as accesses data stored in the memory 120, to perform various functions of the electronic device and process data.
  • the processor 110 may be implemented using at least one of the following hardware forms: digital signal processing (DSP), field-programmable gate array (FPGA), and programmable logic array (PLA).
  • DSP digital signal processing
  • FPGA field-programmable gate array
  • PDA programmable logic array
  • the processor 110 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem.
  • the CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 110 and may be implemented separately via a communication chip.
  • the memory 120 may include random access memory (RAM) or read-only memory (ROM).
  • the memory 120 includes a non-transitory computer-readable storage medium.
  • the memory 120 may be used to store instructions, programs, codes, code sets, or instruction sets.
  • the memory 120 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc.
  • the operating system may be an Android system, including a system deeply developed based on the Android system, an iOS system developed by Apple, including a system deeply developed based on the iOS system, or other systems.
  • the data storage area may also store data created by the electronic device during use, such as a phone book, audio and video data, chat history data, etc.
  • FIG. 13 is a structural diagram of an operating system and user space provided in an embodiment of this specification.
  • the memory 120 can be divided into an operating system space and a user space.
  • the operating system runs in the operating system space, and native and third-party applications run in the user space.
  • the operating system allocates corresponding system resources to different third-party applications.
  • different application scenarios in the same third-party application also have different requirements for system resources. For example, in the local resource loading scenario, the third-party application has higher requirements for disk reading speed; in the animation rendering scenario, the third-party application has higher requirements for GPU performance.
  • the operating system and the third-party application are independent of each other, and the operating system often cannot perceive the current application scenario of the third-party application in a timely manner, resulting in the operating system being unable to perform targeted system resource adaptation according to the specific application scenario of the third-party application.
  • FIG14 is an architectural diagram of the Android operating system in FIG13 provided in an embodiment of this specification. Taking the Android operating system as an example, the programs and data stored in the memory 120 are shown in FIG14 .
  • the memory 120 may store a Linux kernel layer 320, a system runtime library layer 340, an application framework layer 360, and an application layer 380.
  • the Linux kernel layer 320, the system runtime library layer 340, and the application framework layer 360 belong to the operating system space, and the application layer 380 belongs to the user space.
  • the Linux kernel layer 320 provides underlying drivers for various hardware of electronic devices, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, power management, etc.
  • the system runtime library layer 340 provides the main feature support for the Android system through some C/C++ libraries.
  • the SQLite library provides database support
  • the OpenGL/ES library provides 3D drawing support
  • the Webkit library provides browser kernel support.
  • the system runtime layer 340 also provides the Android runtime library (Android runtime), which mainly provides some core libraries that allow developers to use the Java language to write Android applications.
  • the application framework layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management.
  • At least one application runs in the application layer 380. These applications can be native applications that come with the operating system, such as contact applications, SMS applications, clock applications, camera applications, etc.; they can also be third-party applications developed by third-party developers, such as game applications, instant messaging applications, photo beautification applications, etc.
  • FIG. 15 is an architectural diagram of the iOS operating system shown in FIG. 13 , provided in an embodiment of this specification.
  • the iOS system includes: a core operating system layer 420 (Core OS layer), a core services layer 440 (Core Services layer), a media layer 460 (Media layer), and a touchable layer 480 (Cocoa Touch layer).
  • the core operating system layer 420 includes the operating system kernel, drivers, and underlying program frameworks. These underlying program frameworks provide hardware-friendly functionality for use by the program frameworks in the core services layer 440.
  • the core services layer 440 provides application programs with required system services and/or program frameworks, such as the Foundation framework, account framework, advertising framework, data storage framework, network connection framework, geolocation framework, motion framework, and so on.
  • the media layer 460 provides application programs with audio-visual interfaces, such as graphics and image-related interfaces, audio technology-related interfaces, video technology-related interfaces, and an AirPlay interface for wireless audio and video transmission.
  • the touch layer 480 provides various commonly used interface-related frameworks for application development. It is responsible for user touch interaction operations on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit frameworks, and map frameworks.
  • UI message user interface
  • those relevant to most applications include, but are not limited to, the Foundation framework in the core services layer 440 and the UIKit framework in the touchable layer 480.
  • the Foundation framework provides many basic object classes and data types, offering fundamental system services for all applications and having nothing to do with the UI.
  • the classes provided by the UIKit framework are the foundational UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UIs, providing the application infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
  • the input device 130 is used to receive input instructions or data, and the input device 130 includes but is not limited to a keyboard, a mouse, a camera, a microphone or a touch device.
  • the output device 140 is used to output instructions or data, and the output device 140 includes but is not limited to a display device and a speaker.
  • the input device 130 and the output device 140 can be combined, and the input device 130 and the output device 140 are touch screen displays, which are used to receive touch operations on or near the touch screen using any suitable object such as a finger or a touch pen, and to display the user interface of each application.
  • the touch screen display is usually provided on the front panel of the electronic device.
  • the touch screen display can be designed as a full screen, a curved screen or a special-shaped screen.
  • the touch screen display can also be designed as a combination of a full screen and a curved screen, or a combination of a special-shaped screen and a curved screen, which is not limited in this specification.
  • the electronic devices may include more or fewer components than shown, or may combine certain components or arrange the components differently.
  • the electronic devices may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, and other components, which will not be described in detail here.
  • WiFi wireless fidelity
  • the execution entity of each step can be the electronic device described above.
  • the execution entity of each step is the operating system of the electronic device.
  • the operating system can be Android, iOS, or other operating systems, which is not limited in this specification.
  • the electronic device of this specification may also be equipped with a display device.
  • the display device may be any device capable of displaying a display function, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an electronic ink screen, a liquid crystal display (LCD), a plasma display panel (PDP), etc.
  • CR cathode ray tube display
  • LED light-emitting diode display
  • LCD liquid crystal display
  • PDP plasma display panel
  • the electronic device may be a smartphone, a tablet computer, a gaming device, an AR (Augmented Reality) device, a car, a data storage device, an audio player, a video player, a notebook, a desktop computing device, or a wearable device such as an electronic watch, electronic glasses, an electronic helmet, an electronic bracelet, an electronic necklace, or electronic clothing.
  • the processor 110 can be used to call the model training data generation program stored in the memory 120, and specifically perform the following operations: obtain the generative processing model and the target domain data; determine the domain data features corresponding to the target domain data and the model processing labels of the domain data features; update the model processing labels of the domain data features based on the domain data features to obtain the target domain model training data set; and use the target domain model training data set to perform model training on the generative processing model.
  • the processor 110 executes determining specific domain data features from the domain data feature set based on the feature set similarity distribution information, it specifically performs: based on the feature set similarity distribution information, querying the target feature similarity whose feature similarity is less than or equal to the feature similarity threshold from the domain data feature set, determining the domain data feature pairs corresponding to the target feature similarity, and determining the domain data feature pair set based on the domain data feature pairs; determining the feature repetition frequency of each reference domain data feature in the domain data feature pair set; and determining the specific domain data features in the domain data feature set based on the feature repetition frequency of each reference domain data feature.
  • the processor 110 executes model training of the generative processing model using the target domain model training dataset, it specifically executes: performing a first model training process on the generative processing model based on the target domain model training dataset to obtain a first generative processing model; detecting whether the first generative processing model meets the model training end condition; if the first generative processing model does not meet the model training end condition, determining to add new specific domain data features to the domain data feature set, performing label correction on the specific model processing labels carried by the new specific domain data features to obtain the new specific domain data features after label correction, updating the target domain model training dataset based on the new specific domain data features to obtain an updated target domain model training dataset, performing a second model training process on the first generative processing model based on the target domain model training dataset to obtain a second generative processing model, and using the second generative processing model as the first generative processing model to execute in parallel the step of detecting whether the first generative processing model meets the model training end condition; if the first generative processing model meets the model training end condition, determining to
  • the processor 110 determines the domain data features corresponding to the target domain data, it specifically performs: extracting the domain comprehensive features and domain attribute features of the target domain data, extracting the general domain features of the target domain data; determining the domain vertical features based on the domain comprehensive features and domain attribute features, and determining the domain data features based on the domain data features and the general domain data features.
  • F D is the domain data feature
  • fi is the general domain data feature
  • is the third weight coefficient
  • the processor 110 executes the acquisition of the generative processing model and the target domain data, it specifically performs the following: acquiring the generative processing model for the target domain; acquiring the source data and the target domain identifier for the target domain, and identifying the target domain source data from the source data based on the target domain identifier; acquiring the target domain entity definition information, and extracting entity information from the target domain source data based on the target domain entity definition information to obtain the target domain entity information; acquiring the data extraction logic corresponding to the target domain entity information, and extracting data from the target domain entity information based on the data extraction logic to obtain the target domain data.
  • the processor 110 is also suitable for executing: determining the feature similarity between pairwise domain data features in the target domain model training data set; determining a reference feature similarity whose feature similarity is greater than or equal to a similarity threshold and a reference feature pair corresponding to the reference feature similarity; determining two similar domain data features corresponding to the reference feature pair, and performing data fusion processing on the similar domain data features to update the target domain model training data set.
  • the program can be stored in a computer-readable storage medium, and when executed, the program can include the processes in the above-described method embodiments.
  • the storage medium can be a magnetic disk, an optical disk, a read-only memory, or a random access memory.
  • the information including but not limited to user device information, user personal information, etc.
  • data including but not limited to data used for analysis, stored data, displayed data, etc.
  • signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions.
  • the object characteristics, interactive behavior characteristics, and user information involved in this specification are all obtained with full authorization.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • General Engineering & Computer Science (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Software Systems (AREA)
  • Medical Informatics (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

一种模型训练数据生成方法、装置、存储介质及电子设备。该方法包括:获取生成式处理模型和目标领域数据,确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签,基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集,最后采用经过标签更新处理后的目标领域模型训练数据集对生成式处理模型进行模型训练。

Description

模型训练数据生成 技术领域
本说明书涉及数据处理技术领域,尤其涉及用于模型训练数据生成的方法、装置、存储介质及电子设备。
背景技术
近年来,大型语言模型在通用语言处理上表现出色,但在专业垂直领域中表现欠佳,因为这些专业垂直领域涉及复杂的专业知识。为了提升大型语言模型在这些领域中的表现,需要使用特定领域的精准数据进行微调,但获取高质量的专业领域数据并不容易。
发明内容
本说明书提供了一种模型训练数据生成方法、装置、存储介质及电子设备,所述技术方案如下。
第一方面,本说明书提供了一种模型训练数据生成方法,所述方法包括:获取生成式处理模型和目标领域数据;确定所述目标领域数据对应的领域数据特征以及所述领域数据特征的模型处理标签;基于所述领域数据特征对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;采用所述目标领域模型训练数据集对所述生成式处理模型进行模型训练。
第二方面,本说明书提供了一种模型训练数据生成装置,所述装置包括:获取模块,适于获取生成式处理模型和目标领域数据;确定模块,适于确定所述目标领域数据对应的领域数据特征以及所述领域数据特征的模型处理标签;更新模块,适于基于所述领域数据特征对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;训练模块,适于采用所述目标领域模型训练数据集对所述生成式处理模型进行模型训练。
第三方面,本说明书提供一种计算机存储介质,所述计算机存储介质存储有多条指令,所述指令适于由处理器加载并执行上述的方法步骤。
第四方面,本说明书提供一种电子设备,可包括处理器和存储器;其中,所述存储器存储有计算机程序,所述计算机程序适于由所述处理器加载并执行上述的方法步骤。
第五方面,本说明书提供一种计算机程序产品,该计算机程序产品存储有至少一条指令,所述至少一条指令由处理器加载并执行上述任意一项的方法步骤。
本说明书一些实施例提供的技术方案带来的有益效果至少包括:通过确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签,基于领域数据特征确定领域数据特征中与其他领域数据特征差异性较大的领域数据特征,而与其他领域数据特征差异性较大的领域数据特征对应的模型处理标签往往容易打标错误,由此基于领域数据特征对领域数据特征中容易出现打标错误的领域数据特征的模型处理标签进行标签更新,从而得到目标领域模型训练数据集,提高了目标领域数据特征集的训练样本质量,最后采用得到的目标领域模型训练数据集对生成式处理模型进行模型训练,从而得到针对目标领域的高质量处理模型。
附图说明
图1为本说明书提供的一种模型训练数据生成系统的场景示意图。
图2为本说明书实施例提供了一种模型训练数据生成方法的流程示意图。
图3为本说明书实施例提供了一种确定目标领域模型训练数据集的流程示意图。
图4为本说明书实施例提供了一种确定特异性领域数据特征的流程示意图。
图5为本说明书实施例提供了又一种确定特异性领域数据特征的流程示意图。
图6为本说明书实施例提供了另一种确定特异性领域数据特征的流程示意图。
图7为本说明书实施例提供了一种确定领域数据特征集中新增特异性领域数据特征的流程示意图。
图8为本说明书实施例提供了一种确定目标领域数据对应的领域数据特征的流程示意图。
图9为本说明书实施例提供了一种生成式处理模型和目标领域数据的流程示意图。
图10为本说明书实施例提供的一种更新目标领域模型训练数据集的流程示意图。
图11为本说明书实施例提供的一种模型训练数据生成装置。
图12为本说明书实施例提供的电子设备的结构方框图。
图13为本说明书实施例提供的一种操作系统和用户空间的结构示意图。
图14为本说明书实施例提供的一种图13中安卓操作系统的架构图。
图15为本说明书实施例提供的一种图13中IOS操作系统的架构图。
具体实施方式
下面将结合本说明书中的附图,对本说明书中的技术方案进行清楚、完整地描述,显然,所描述的实施例仅仅是本说明书一部分实施例,而不是全部的实施例。基于本说明书中的实施例,本领域普通技术人员在没有做出创造性劳动前提下所获得的所有其他实施例,都属于本说明书保护的范围。
在本说明书的描述中,需要理解的是,术语“第一”、“第二”等仅用于描述目的,而不能理解为指示或暗示相对重要性。在本说明书的描述中,需要说明的是,除非另有明确的规定和限定,“包括”和“具有”以及它们任何变形,意图在于覆盖不排他的包含。例如包含了一系列步骤或单元的过程、方法、系统、产品或设备没有限定于已列出的步骤或单元,而是可选地还包括没有列出的步骤或单元,或可选地还包括对于这些过程、方法、产品或设备固有的其他步骤或单元。对于本领域的普通技术人员而言,可以具体情况理解上述术语在本说明书中的具体含义。此外,在本说明书的描述中,除非另有说明,“多个”是指两个或两个以上。“和/或”,描述关联对象的关联关系,表示可以存在三种关系,例如,A和/或B,可以表示:单独存在A,同时存在A和B,单独存在B这三种情况。字符“/”一般表示前后关联对象是一种“或”的关系。
下面结合具体的实施例对本说明书进行详细说明。
请参见图1,图1为本说明书提供的一种模型训练数据生成系统的场景示意图。如图1所示,所述模型训练数据生成系统至少可以包括客户端集群和服务平台100。
所述客户端集群可以包括至少一个客户端,如图1所示,具体包括用户1对应的客户端1、用户2对应的客户端2、…、用户n对应的客户端n,n为大于0的整数。
客户端集群中的各客户端可以是具备通信功能的电子设备,该电子设备包括但不限于:可穿戴设备、手持设备、个人电脑、平板电脑、车载设备、智能手机、计算设备或连接到无线调制解调器的其它处理设备等。在不同的网络中电子设备可以叫做不同的名称,例如:用户设备、接入终端、用户单元、用户站、移动站、移动台、远方站、远程终端、移动设备、用户终端、终端、无线通信设备、用户代理或用户装置、蜂窝电话、无绳电话、个人数字处理(personal digital assistant,PDA)、5G网络或未来演进网络中的电子设备等。
所述服务平台100可以是单独的服务器设备,例如:机架式、刀片、塔式、或者机柜式的服务器设备,或采用工作站、大型计算机等具备较强计算能力硬件设备;也可以是采用多个服务器组成的服务器集群,所述服务集群中的各服务器可以是以对称方式组成的,其中每台服务器在事务链路中功能等价、地位等价,各服务器均可单独对外提供服务,所述单独提供服务可以理解为无需另外的服务器的辅助。
在本说明书的一个或多个实施例中,服务平台100与客户端集群中的至少一个客户端可建立通信连接,基于该通信连接完成模型训练数据生成过程中数据的交互,诸如线上事务数据交互,如服务平台100可基于本说明书的模型训练数据生成方法得到的目标领域模型训练数据集可实现对生成式处理模型进行模型训练。
服务平台100与客户端集群中的至少一个客户端通过网络建立通信连接进行交互通信,其中,网络可以是无线网络,也可以是有线网络,无线网络包括但不限于蜂窝网络、无线局域网、红外网络或蓝牙网络,有线网络包括但不限于以太网、通用串行总线(universal serial bus,USB)或控制器局域网络。在说明书一个或多个实施例中,使用包括超文本标记语言(Hyper Text Mark-up Language,HTML)、可扩展标记语言(Extensible Markup Language,XML)等的技术和/或格式来代表通过网络交换的数据(如目标压缩包)。此外还可以使用诸如安全套接字层(Secure Socket Layer,SSL)、传输层安全(Transport Layer Security,TLS)、虚拟专用网络(Virtual Private Network,VPN)、网际协议安全(Internet Protocol Security,IPsec)等常规加密技术来加密所有或者一些链路。在另一些实施例中,还可以使用定制和/或专用数据通信技术取代或者补充上述数据通信技术。
本说明书所提供的模型训练数据生成系统实施例与一个或多个实施例中的所述模型训练数据生成方法属于同一构思,在说明书一个或多个实施例涉及的所述模型训练数据生成方法对应的执行主体可以是上述服务平台100;在说明书一个或多个实施例涉及的所述模型训练数据生成方法对应的执行主体也可以是客户端所对应的电子设备,具体基于实际应用环境确定。模型训练数据生成系统实施例其体现实现过程可详见下述的方法实施例,这里不再赘述。
基于图1所示的场景示意图,下面对本说明书一个或多个实施例提供的模型训练数据生成方法进行详细介绍。
请参见图2,图2为本说明书实施例提供了一种模型训练数据生成方法的流程示意图,该方法可依赖于计算机程序实现,可运行于基于冯诺依曼体系的模型训练数据生成装置上。该计算机程序可集成在应用中,也可作为独立的工具类应用运行。所述模型训练数据生成装置可以为服务平台。
具体的,该模型训练数据生成方法包括以下步骤。
S202:获取生成式处理模型和目标领域数据。
其中,生成式处理模型可以为生成式大语言模型,生成式大语言模型主要使用深度学习技术如神经网络构建而成,这类模型能够基于大量文本数据进行学习,并结合复杂的算法来理解、解释和生成人类语言。
这里,生成式处理模型可以用于应用于目标领域,以进行目标领域的数据处理。一般的,生成式处理模型可以较好的处理在通用语言处理任务,但是在专业垂直领域的执行能力却有所不足,因为这些领域往往充满了复杂的专业术语和独特的操作规则。
目标领域指的是特定问题领域或知识领域,它可以是一组领域相关的专业术语、概念、规则、流程和实体集合的集合体。目标领域可以是任何具体的知识领域,如金融、医疗、教育、电子商务等,以及任何具体的问题领域,具体可以为应用领域,如自然语言处理、图像识别、机器学习等。
因此,除了获取针对目标领域的生成式处理模型外,还需要获取针对目标领域的目标领域数据,目标领域数据可以是各种类型的数据,如结构化数据、非结构化数据、半结构化数据等。结构化数据指的是具有明确格式和规则的数据,如数据库中的表格数据、关系型数据等;非结构化数据则是指没有明确格式和规则的数据,如文本数据、图像数据、音频数据等;半结构化数据则是介于结构化数据和非结构化数据之间的数据,如XML(可扩展标记语言,Extensible Markup Language)文档、JSON(JS对象简谱,JavaScript Object Notation)数据等。
同时,目标领域数据还可以为金融领域、医疗领域、教育领域或交通领域等。以金融领域为例,金融数据可以包括证券市场数据、汇率数据、金融机构数据、金融产品数据等。医疗领域的数据则可以包括患者数据、药品数据、医疗设备数据、疾病诊断数据等。不同领域的数据具有各自的特点和应用场景,因此在进行数据处理和分析时需要考虑领域的特殊性。
S204:确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签。
其中,在获得目标领域数据后,通过对目标领域数据进行数据特征提取,获得领域数据特征。这里,领域数据特征兼具通用领域维度的特征和领域垂直维度的特征。
应理解,领域数据特征的呈现形式包括但不限于领域数据特征向量、领域数据特征矩阵、领域数据特征图、领域数据特征值等。这些呈现形式的选择取决于具体应用场景。
领域数据特征向量是一种简洁的领域数据特征表示方式,可表示数据的属性或特征。例如,在文本分类任务中,领域数据特征向量可表示文本中的关键词、短语或主题。领域数据特征矩阵则是一种更为详细的领域数据特征表示方式。领域数据特征图则是一种直观的领域数据特征表示方式,通常用于可视化领域数据特征之间的关系和网络结构。
具体的,在对目标领域数据进行数据特征提取时,可以从目标领域数据的通用领域维度和目标领域对应的领域垂直维度对目标领域数据进行两个维度的数据特征的提取,之后将这两个维度的数据特征进行特征融合得到领域数据特征。当领域数据特征为领域数据特征向量时,数据特征的提取包括数据向量化过程;当领域数据特征为领域数据特征矩阵时,数据特征的提取包括数据矩阵化过程;当领域数据特征为领域数据特征图时,数据特征的提取包括数据可视化过程;当领域数据特征为领域数据特征值时,数据特征的提取包括数值化过程。
由于领域数据特征兼具通用领域维度的特征和领域垂直维度的特征,因此生成式处理模型在学习通用领域维度的特征后,可以使得模型可以更加高效精准的处理通用处理任务;并且,生成式处理模型在学习领域垂直维度的特征后,可以使模型可以在学习通用领域维度的特征基础上更加高效精准的处理目标领域的处理任务。
为了基于领域数据特征对生成式处理模型进行模型训练,因此可以针对生成式处理模型处理的任务类型对领域数据特征进行自动化标注处理,如生成式处理模型处理的任务类型可以为识别目标领域的文本的真或假。此时,对领域数据特征进行自动化标注处理结果可以为将各领域数据特征基于标签标注规则或者预训练模型自动标注真或假的模型处理标签。
其中,确定领域数据特征的模型处理标签可包括:获取数据标注规则集合,基于数据标注规则集合对领域数据特征进行数据标注,得到领域数据特征对应的模型处理标签。
数据标注规则集合可以基于专家经验以及目标领域的专业知识库等确定。
可选地,确定领域数据特征的模型处理标签,包括:基于预训练模型对领域数据特征进行标注处理,得到领域数据特征对应的模型处理标签。
这里,预训练模型可以为预先训练完成的可对领域数据特征进行标注的模型。当然,在其他的实施例中,还可以基于预训练模型和数据标注规则集合进行自动化标注处理,具体的,自动化标注处理的过程可以用以下公式表示:
Lpre(D)=AutoAnnotate(D;R,M)。
其中,Lpre(D)表示领域数据特征D的模型处理标签;R表示数据标注规则集合;M表示预训练模型;AutoAnnotate表示自动化标注函数。
S206:基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集。
在确定了领域数据特征后,基于领域数据特征之间的特征相似性确定领域数据特征中与其他领域数据特征差异性较大的领域数据特征,而与其他领域数据特征差异性较大的领域数据特征对应的模型处理标签往往容易打标错误。由此,可基于领域数据特征对领域数据特征中容易出现打标错误的领域数据特征的模型处理标签进行标签更新。
因此,可以对与其他领域数据特征差异性较大的领域数据特征对应的模型处理标签进行标签更新,从而实现基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到经过更新后的目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签,之后,基于目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签得到目标领域模型训练数据集。
S208:采用目标领域模型训练数据集对生成式处理模型进行模型训练。
其中,在得到目标领域模型训练数据集后,利用目标领域模型训练数据集对生成式处理模型进行高质量的模型训练,从而使得生成式处理模型可以准确学习针对目标领域的目标领域数据对应的专业领域特征。
在本说明书提供的实施例中,通过确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签,基于领域数据特征确定领域数据特征中与其他领域数据特征差异性较大的领域数据特征,而与其他领域数据特征差异性较大的领域数据特征对应的模型处理标签往往容易打标错误,由此基于领域数据特征对领域数据特征中容易出现打标错误的领域数据特征的模型处理标签进行标签更新,从而得到目标领域模型训练数据集,提高了目标领域数据特征集的训练样本质量,最后采用得到的目标领域模型训练数据集对生成式处理模型进行模型训练,从而得到针对目标领域的高质量处理模型。
请参见图3,图3为本说明书实施例提供了一种确定目标领域模型训练数据集的流程示意图。如图3所示,S206中基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集,包括以下步骤。
S302:基于领域数据特征确定特异性领域数据特征,对特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签。
其中,在得到具有模型处理标签的领域数据特征后,由于模型处理标签通常为自动化标注处理得到,其准确性难以得到保证。因此,需要对领域数据特征的模型处理标签进行标签修正处理,然而由于对领域数据特征的模型处理标签进行标签修正处理涉及大量目标领域的专业知识。因此,为了提高标签修正处理效率和节约成本,可以对领域数据特征进行选择性标签修正处理。
因此,可以基于领域数据特征之间的特征相似性确定领域数据特征中与其他领域数据特征差异性较大的领域数据特征,即特异性领域数据特征,而与其他领域数据特征差异性较大的特异性领域数据特征对应的特异性模型处理标签往往容易打标错误。
因此,可以对特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签。这里,可以基于训练模型进行标签修正处理;或,专家介入进行标签修正处理。
S304:基于修正模型处理标签对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集。
其中,确定目标领域数据对应的领域数据特征中与特异性领域数据特征对应的特征,之后基于特异性领域数据特征对应的修正模型处理标签对与特异性领域数据特征对应的特征的标签进行标签更新处理,即进行标签替换,从而得到目标领域数据对应的领域数据特征,以及更新后目标领域数据对应的领域数据特征的模型处理标签,基于目标领域数据对应的领域数据特征,以及更新后目标领域数据对应的领域数据特征的模型处理标签得到目标领域模型训练数据集。此时的目标领域数据特征集在针对目标领域的领域专业性和领域特征的准确性均具有较高的保证。
在本说明书提供的实施例中,基于领域数据特征确定领域数据特征中与其他领域数据特征差异性较大的领域数据特征,即特异性领域数据特征,相较于普通的领域数据特征,特异性领域数据特征在打标签时更容易出现错误,因此需要对特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签,之后通过特异性领域数据特征对应的修正模型处理标签对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集,从而有效提高了目标领域数据特征集的训练样本质量,避免了人工专家的介入,节省了成本,提高了目标领域数据特征集的生成质量以及效率。
在本说明书提供的一种实施例中,S302中基于领域数据特征确定特异性领域数据特征,包括:对领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于领域数据特征集确定特异性领域数据特征。
对领域数据特征进行选择性标签修正处理过程中,首先对领域数据特征进行聚类处理,如Cluster(D;K),K为聚类数,K大于或等于2,D为领域数据特征,Cluster(D;K)是根据聚类数K对D进行的聚类函数。这里,可以基于领域数据特征彼此之间的整体相似度进行聚类处理;或,基于领域数据特征彼此之间的一个或多个维度的相似度进行聚类处理。
聚类处理后得到至少两类领域数据特征集,每个领域数据特征集中的领域数据特征在一个或多个维度上具有相似性。因此,可以基于领域数据特征集中的领域数据特征之间的相似性程度来确定特异性领域数据特征,特异性领域数据特征可以为领域数据特征集中与其它领域数据特征相似度较低的特征。换言之,特异性领域数据特征可以理解为对生成式处理模型性能提升较为关键的特征。
由于特异性领域数据特征在领域数据特征集中与其它领域数据特征相似度较低,因此,特异性领域数据特征的特异性模型处理标签存在误差的可能性较高,此时可以对每个领域数据特征集中的特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,从而得到针对特异性领域数据特征的修正模型处理标签。
在本说明书提供的实施例中,由于模型处理标签往往为自动化标注处理得到,准确性难以得到保证,因此需要对领域数据特征的模型处理标签进行标签修正处理。标签修正处理之前首先进行聚类处理,由于聚类处理后得到的每个领域数据特征集中的领域数据特征在一个或多个维度上具有相似性,而特异性领域数据特征在领域数据特征集中与其它领域数据特征相似度较低。
因此,特异性领域数据特征的特异性模型处理标签存在误差的可能性较高,此时可对每个领域数据特征集中的特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,提高目标领域数据特征集的训练样本质量,最后采用得到的目标领域模型训练数据集对生成式处理模型进行模型训练,得到针对目标领域的高质量处理模型。
请参见图4,图4为本说明书实施例提供了一种确定特异性领域数据特征的流程示意图。具体的:上述实施例中对领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于领域数据特征集确定特异性领域数据特征,包括以下步骤。
S402:计算两两领域数据特征之间的特征相似度,基于特征相似度进行聚类,得到至少两类领域数据特征集。
其中,可以计算两两领域数据特征之间的特征相似度,如基于相似度计算公式进行计算或欧氏距离相似度计算公式来计算特征相似度。在得到两两领域数据特征之间的特征相似度后,基于特征相似度对领域数据特征进行划分,将特征相似度较高的领域数据特征划分在相同的领域数据特征集,得到至少两类领域数据特征集。
在其他的实施例中,还可以确定领域数据特征对应的数据点坐标,如领域数据特征为(1,1,1,1)时,此时的数据点坐标可以为(1,1,1,1)。在得到各领域数据特征对应的数据点坐标后,确定各领域数据特征对应的数据点在空间中的分布情况,之后将在空间中聚集在一起的数据点对应的领域数据特征划分在相同的领域数据特征集,最终得到至少两类领域数据特征集。
S404:基于领域数据特征集的所有特征相似度,确定特征集相似度分布信息。
其中,获取各领域数据特征集的所有特征相似度,即各领域数据特征集中两两领域数据特征之间的特征相似度,之后基于得到的全部特征相似度进行统计分析处理,得到特征集相似度分布信息。
具体的,特征集相似度分布信息可以为各特征相似度区间分布的领域数据特征对。如,特征相似度区间为40%~50%之间分布的领域数据特征对及其对应的特征相似度,特征相似度区间为70%~80%之间分布的领域数据特征对及其对应的特征相似度。
S406:基于特征集相似度分布信息从领域数据特征集中确定特异性领域数据特征。
特征集相似度分布信息记录了各领域数据特征在各特征相似度区间的出现频次。
容易理解的,在低特征相似度区中出现频次较高的领域数据特征可以认为是与其它领域数据特征相似度较低的向量,因此可以作为特异性领域数据特征。
在本说明书提供的实施例中,通过两两领域数据特征之间的特征相似度进行聚类,从而得到至少两类领域数据特征集,之后确定各领域数据特征集中的特征集相似度分布信息,从而通过特征集相似度分布信息快速确定与其它领域数据特征相似度较低的特异性领域数据特征。
请参见图5,图5为本说明书实施例提供了又一种确定特异性领域数据特征的流程示意图。具体的:S406中基于特征集相似度分布信息从领域数据特征集中确定特异性领域数据特征,包括以下步骤。
S502:基于特征集相似度分布信息,从领域数据特征集中查询特征相似度小于或等于特征相似度阈值的目标特征相似度,确定目标特征相似度对应的领域数据特征对,基于领域数据特征对确定领域数据特征对集合。
其中,特征相似度阈值可以基于特征集相似度分布信息对应进行灵活设置,如特征集相似度分布信息中特征相似度区间最小为30%~40%,则特征相似度阈值可以设置为40%或50%等。当然,特征相似度阈值也可以人工预先设置。不同的领域数据特征集对应的特征相似度阈值可以相同或不同。
在确定特征相似度阈值后,从领域数据特征集中查找特征集相似度分布信息中特征相似度小于或等于特征相似度阈值的目标特征相似度,每一个目标特征相似度对应着一对领域数据特征,即领域数据特征对,将各目标特征相似度对应的领域数据特征对放入同一个集合中,得到该领域数据特征集对应的领域数据特征对集合。
S504:确定领域数据特征对集合中各参考领域数据特征的特征重复频次。
其中,领域数据特征对集合中各领域数据特征对的特征相似度小于或等于特征相似度阈值,因此,领域数据特征对集合中各领域数据特征对的向量相似性较低,此时,可以统计领域数据特征对集合中各参考领域数据特征的特征重复频次。
当参考领域数据特征的特征重复频次越高,则表明该参考领域数据特征与其他的参考领域数据特征的区别较大;当参考领域数据特征的特征重复频次越低,则表明该参考领域数据特征与其他的参考领域数据特征的区别较小。
S506:基于各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
其中,当参考领域数据特征的特征重复频次越高,则该参考领域数据特征与其他的参考领域数据特征的区别较大,此时该参考领域数据特征可以被认为是领域数据特征集中的特异性领域数据特征。这里,特异性领域数据特征可以为一个或多个。
具体的,可以对各参考领域数据特征的特征重复频次进行排序处理,如降序排列,将排序在前的特征重复频次对应的参考领域数据特征作为领域数据特征集中的特异性领域数据特征。
本说明书提供的实施例中,通过特征相似度阈值滤除领域数据特征集中相似度较高的领域数据特征对,得到领域数据特征集中相似度较低的领域数据特征对,即领域数据特征对集合,之后统计领域数据特征对集合中各参考领域数据特征的特征重复频次,当参考领域数据特征的特征重复频次越高,则表明该参考领域数据特征与其他的参考领域数据特征的区别较大,此时该参考领域数据特征可以被认为是领域数据特征集中的特异性领域数据特征。
请参见图6,图6为本说明书实施例提供了另一种确定特异性领域数据特征的流程示意图。具体的:S506中基于各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征,包括以下步骤。
S602:获取模型训练指标与标签修正比例的指标比例映射关系,确定针对生成式处理模型的目标模型训练指标,基于目标模型训练指标和指标比例映射关系确定目标标签修正比例。
其中,模型训练指标为生成式处理模型训练后的预期性能指标,不同的标签修正比例对应不同的模型训练指标,当标签修正比例越高时,生成式处理模型训练后的性能指标也越高;标签修正比例越低时,生成式处理模型训练后的性能指标也越低。标签修正比例为领域数据特征集中进行标签修正处理的特异性领域数据特征的比例。具体的,目标标签修正比例可以表示为Stratify(D;P),P为目标模型训练指标,Stratify可以为指标比例映射关系函数。
在得到指标比例映射关系后,获取生成式处理模型的目标模型训练指标,如处理目标领域的目标领域任务时,正确率为95%。通过指标比例映射关系可以确定目标模型训练指标对应的目标标签修正比例。
S604:基于目标标签修正比例和领域数据特征集的领域数据特征数量,确定特异性领域数据特征数量。
其中,在得到目标标签修正比例后,将目标标签修正比例和领域数据特征集的领域数据特征数量相乘,可以计算得到需要进行标签修正处理的特征的数量,即特异性领域数据特征数量。
S606:基于特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
其中,步骤S404中已经获得了各参考领域数据特征的特征重复频次,可以将该特征重复频次进行降序排列。当特异性领域数据特征数量为X时,可以将降序排列中前X个特征重复频次对应的参考领域数据特征作为特异性领域数据特征,从而得到领域数据特征集中的特异性领域数据特征。特异性领域数据特征可以表示为Cluster(D;K)和Stratify(D;P)的交集。
本说明书实施例通过目标模型训练指标和指标比例映射关系确定目标标签修正比例,从而确定特异性领域数据特征数量,进而基于各参考领域数据特征的特征重复频次参考特异性领域数据特征数量进行特异性领域数据特征筛选,保证对生成式处理模型训练质量的同时,减少了进行标签修正处理的工作量,提升了模型训练数据生成效率。
在本说明书提供的一种实施例中,S208中采用目标领域模型训练数据集对生成式处理模型进行模型训练,包括:基于目标领域模型训练数据集对生成式处理模型进行第一模型训练处理,得到第一生成式处理模型;检测第一生成式处理模型是否满足模型结束训练条件;若第一生成式处理模型不满足模型结束训练条件,则确定领域数据特征集中新增特异性领域数据特征,对新增特异性领域数据特征携带的特异性模型处理标签进行标签修正得到标签修正后的新增特异性领域数据特征,基于新增特异性领域数据特征对目标领域模型训练数据集进行更新得到更新后的目标领域模型训练数据集,基于目标领域模型训练数据集对第一生成式处理模型进行第二模型训练处理,得到第二生成式处理模型,将第二生成式处理模型作为第一生成式处理模型并行执行检测第一生成式处理模型是否满足模型结束训练条件的步骤;若第一生成式处理模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
在得到目标领域模型训练数据集后,可以采用目标领域模型训练数据集对生成式处理模型进行第一模型训练处理,得到第一生成式处理模型。第一生成式处理模型可能没有达到模型结束训练条件,因此可以再次对目标领域模型训练数据集中除特异性领域数据特征外的向量部分进行标签修正处理,从而再次对第一生成式处理模型进行训练,之后重复进行,即,利用每轮标注数据对上一轮得到的模型进行增量训练,直到得到的模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
具体的,可以用如下表达式表达:
Mi+1=Train(Mi,Si)。
其中,Mi+1是第i+1轮迭代后的模型,i为大于或等于0的整数;Mi是第i轮的模型;Si是第i轮更新后的目标领域模型训练数据集;Train是进行型训练处理。
通过上述策略的实施,本方案不仅显著减轻了标签修正处理的标注负担,同时也通过多轮模型训练提高了领域数据标注的精确度,优化了目标领域模型训练数据集的质量和可用性,为领域特定的模型训练奠定了良好的基础,使得训练得到的目标生成式处理模型可以精准高质量的处理目标领域的目标领域任务。同时,通过持续的迭代循环,利用新确定的标注数据不断完善模型,逐步减少对高质量标注的需求并优化整个标注过程。
请参见图7,图7为本说明书实施例提供了一种确定领域数据特征集中新增特异性领域数据特征的流程示意图。上述实施例中确定领域数据特征集中新增特异性领域数据特征,包括以下步骤。
S702:确定第一生成式处理模型的当前模型训练指标,基于当前模型训练指标与目标模型训练指标的指标差值确定新增特异性领域数据特征数量。
其中,第一生成式处理模型的当前模型训练指标为进行训练后得到的第一生成式处理模型的训练效果指标。
一般的,当前模型训练指标与目标模型训练指标会存在一定的差值,当当前模型训练指标小于目标模型训练指标时,则可以基于当前模型训练指标与目标模型训练指标的指标差值确定新增特异性领域数据特征数量;当当前模型训练指标大于目标模型训练指标时,表明模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
指标差值确定新增特异性领域数据特征数量可以基于预先建立的映射关系或者预先确定的模型来确定。
S704:基于新增特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中新增特异性领域数据特征。
其中,在确定新增特异性领域数据特征数量后,可以再次基于各参考领域数据特征的特征重复频次,从未被确定为特异性领域数据特征的特征重复频次中再次选择新增特异性领域数据特征数量的新增特异性领域数据特征。
这里,新增特异性领域数据特征选择的方式与步骤S506类似,因此不再赘述。
本说明书实施例通过当前模型训练指标与目标模型训练指标的指标差值确定新增特异性领域数据特征数量,从而通过新增特异性领域数据特征数量和各参考领域数据特征的特征重复频次确定领域数据特征集中新增特异性领域数据特征,同样保证对生成式处理模型训练质量的同时,减少了进行标签修正处理的工作量,提升了模型训练数据生成效率。
请参见图8,图8为本说明书实施例提供了一种确定目标领域数据对应的领域数据特征的流程示意图。S204中确定目标领域数据对应的领域数据特征,包括以下步骤。
S802:提取目标领域数据的领域综合特征和领域属性特征,提取目标领域数据的通用领域特征。
其中,从目标领域数据的通用领域维度和目标领域对应的领域垂直维度对目标领域数据进行两个维度的数据特征的提取,从而得到目标领域对应的领域垂直维度对应的领域综合特征和领域属性特征,以及通用领域维度对应的通用领域特征。
领域综合特征综合了各领域属性特征,领域属性特征可以理解为对目标领域进行再次划分得到的各子领域的属性特征,领域综合特征从宏观角度对目标领域的特征进行了汇总,领域属性特征从具体的子领域的属性特征对目标领域的特征进行了拆分。
S804:基于领域综合特征和领域属性特征确定领域垂类特征,基于领域数据特征和通用领域数据特征确定领域数据特征。
其中,由于领域综合特征从宏观角度对目标领域的特征进行了汇总,领域属性特征从具体的子领域的属性特征对目标领域的特征进行了拆分,因此,基于领域综合特征和领域属性特征可以从宏观和具体细分子领域两个角度综合确定领域垂类特征,可以有效避免领域特征的丢失。
领域数据特征基于领域数据特征和通用领域数据特征确定,使得生成式处理模型在学习通用领域数据特征后,可以使得模型可以更加高效精准的处理通用处理任务;并且,生成式处理模型在学习领域数据特征后,可以使模型可以在学习通用领域数据特征基础上更加高效精准的处理目标领域的处理任务。
具体的,S804可以包括以下步骤:确定领域综合特征对应的第一权重系数,确定领域属性特征对应的第二权重系数;将领域综合特征、领域属性特征、第一权重系数、第二权重系数采用第一计算式得到领域垂类特征;第一计算式满足以下公式:
FL=α*fc+β*fb
其中,FL是领域垂类特征,α是第一权重系数,β是第二权重系数,fc是领域综合特征,fb是领域属性特征;确定通用领域数据特征对应的第三权重系数,第一权重系数、第二权重系数和第三权重系数的和为1;将领域垂类特征、通用领域数据特征、第三权重系数采用第二计算式得到领域数据特征,基于领域数据特征确定领域数据特征;第二计算式满足以下公式:
FD=FL+γ*fi
其中,FD是领域数据特征,fi是通用领域数据特征,γ是第三权重系数。
应理解,第一权重系数、第二权重系数和第三权重系数可以基于领域综合特征、领域属性特征以及通用领域数据特征的特征重要性进行确定。
请参见图9,图9为本说明书实施例提供了一种生成式处理模型和目标领域数据的流程示意图。具体的:S202中获取生成式处理模型和目标领域数据,包括以下步骤。
S902:获取针对目标领域的生成式处理模型。
其中,S902中获取针对目标领域的生成式处理模型可以参考S202中的相关描述,这里不再赘述。
S904:获取源数据和针对目标领域的目标领域标识符,基于目标领域标识符从源数据中识别目标领域源数据。
其中,目标领域标识符可为目标领域中的领域关键词,基于领域关键词可以识别对应的目标领域的数据。因此,可基于目标领域标识符从源数据中识别目标领域源数据。
S906:获取目标领域实体定义信息,基于目标领域实体定义信息从目标领域源数据中进行实体信息提取,得到目标领域实体信息。
其中,目标领域实体定义信息可以为目标领域中专业名词或专业术语等对应的定义信息,通过目标领域实体定义信息从目标领域源数据中进行实体信息提取,得到目标领域实体信息,目标领域实体信息可以包括专业名词或专业术语等的描述语句。
S908:获取目标领域实体信息对应的数据提取逻辑,基于数据提取逻辑对目标领域实体信息进行数据提取,得到目标领域数据。
其中,目标领域实体信息包括专业名词或专业术语等的描述语句。因此,不同的目标领域实体信息中对应的描述存在区别,可以通过获取目标领域实体信息对应的数据提取逻辑对目标领域实体信息进行数据特征提取,从而得到目标领域数据。
请参见图10,图10为本说明书实施例提供的一种更新目标领域模型训练数据集的流程示意图。如图9所示,该方法包括以下步骤。
S1002:确定目标领域模型训练数据集中两两领域数据特征之间的特征相似度。
其中,特征相似度计算公式可以为
其中,xi和yi分别代表目标领域模型训练数据集中领域数据特征对应的两个向量。两两向量之间的向量相似度越高,表明这两个特征的特征相似性越强。即这两个特征的关联性越强。
S1004:确定特征相似度大于或等于相似度阈值的参考特征相似度以及参考特征相似度对应的参考特征对。
其中,相似度阈值可以人为预设,如相似度阈值可以设置为98%等。大于或等于相似度阈值的参考特征对高度相似,为了节省数据存储空间,可以对参考特征对对应的两个特征进行数据融合处理。
S1006:确定参考特征对对应的两个相似领域数据特征,对相似领域数据特征进行数据融合处理,以更新目标领域模型训练数据集。
其中,可以将目标领域模型训练数据集中参考特征对对应的两个相似领域数据特征相同的部分进行提取,将不同的部分进行特征拼接,从而实现对对应的两个相似领域数据特征进行数据融合处理,以更新目标领域模型训练数据集。
此外,在得到目标领域模型训练数据集后,目标领域模型训练数据集可被存储在专为高效检索和计算优化的向量数据库中。此向量数据库采用复杂的索引机制,如近似最邻搜索算法,以及高级数据表示技术,便于在多个维度上进行查询和分析,诸如唯一数据标识符、特征相似度和高斯概率分布等。同时可利用高斯概率分布来估计数据的分布特性,当分布特性不满足要求时,可调整参数对目标领域模型训练数据集进行重新生成。
在本说明书提供的一种具体的实施例中,模型训练数据生成方法具体可以包括四个阶段,数据引入阶段、数据处理阶段、数据标注阶段以及领域知识库阶段,其中,在数据引入阶段对源数据进行领域数据管理,源数据可以包括文本、图表、音视频等,领域数据管理可以为对源数据进行目标领域标识符确定、目标领域实体定义信息确定、数据提取逻辑确定、检验数据提取逻辑,之后获取目标领域数据,在目标领域数据的提取过程中提取途径可以包括网络爬虫、光学字符识别、直接获取文档信息等。依据数据提取逻辑执行高精度的数据摘取;最后,所采集的数据将根据数据验证规则进行严格的质量控制,从而保障了所引入数据在准确性和完整性上的无误性。
在数据处理阶段,将得到的目标领域数据先通过标准处理流程进行处理,之后通过领域自定义流程进行处理,最后借助通用大模型对数据中的通用知识成分进行智能修正。这一机制显著减少了对人工标注的依赖,从而降低了标注成本并提升了数据处理的效率。标准处理流程包括数据切分、数据清洗、标准格式化以及数据增强。
领域自定义流程包括领域分类、通用维度处理、领域属性处理以及属性增强,考虑到不同领域数据的处理需求差异性,本方案构建了一个松耦合、高度可扩展的组件式数据处理架构。该架构支持用户自定义领域处理流程,这包括但不限于领域分类算法、多维度通用处理、属性特定处理以及属性增广等多方面的功能模块。最后通过通用大模型进行数据输出,将输出的数据进行向量化处理,得到领域数据特征对应的领域数据特征向量。
在数据标注阶段,包括自动化标注处理、标签修正处理以及循环对新增特异性领域数据特征进行标签修正处理。
在领域知识库阶段,可以对目标领域模型训练数据集进行存储,领域知识库阶段包括计算数据相似度、计算数据概率分布信息、目标领域模型训练数据集以及其中对应的模型处理标签等。
下面将结合图11,图11为本说明书实施例提供的一种模型训练数据生成装置,下面对本说明书提供的模型训练数据生成装置进行详细介绍。需要说明的是,图11所示的模型训练数据生成装置,用于执行本说明书图1~图10所示实施例的方法,为了便于说明,仅示出了与本说明书相关的部分,具体技术细节未揭示的,请参照本说明书图1~图10所示的实施例。
请参见图11,其示出本说明书的模型训练数据生成装置的结构示意图。该模型训练数据生成装置1可以通过软件、硬件或者两者的结合实现成为用户终端的全部或一部分。根据一些实施例,该模型训练数据生成装置1包括获取模块11、确定模块12、更新模块13和训练模块14,其中:获取模块11,适于获取生成式处理模型和目标领域数据;确定模块12,适于确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签;更新模块13,适于基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;训练模块14,适于采用目标领域模型训练数据集对生成式处理模型进行模型训练。
可选地,更新模块13包括:修正单元,适于基于领域数据特征确定特异性领域数据特征,对特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签;更新单元,适于基于修正模型处理标签对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集。
可选地,修正单元还适于对领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于领域数据特征集确定特异性领域数据特征。
可选地,修正单元包括:计算子单元,适于计算两两领域数据特征之间的特征相似度,基于特征相似度进行聚类,得到至少两类领域数据特征集;第一确定子单元,适于基于领域数据特征集的所有特征相似度,确定特征集相似度分布信息;第二确定子单元,适于基于特征集相似度分布信息从领域数据特征集中确定特异性领域数据特征,基于特异性领域数据特征确定特异性领域数据特征。
可选地,第二确定子单元包括:领域数据特征对集合确定子单元,适于基于特征集相似度分布信息,从领域数据特征集中查询特征相似度小于或等于特征相似度阈值的目标特征相似度,确定目标特征相似度对应的领域数据特征对,基于领域数据特征对确定领域数据特征对集合;特征重复频次确定子单元,适于确定领域数据特征对集合中各参考领域数据特征的特征重复频次;领域数据特征确定子单元,适于基于各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
可选地,领域数据特征确定子单元包括:获取子单元,适于获取模型训练指标与标签修正比例的指标比例映射关系,确定针对生成式处理模型的目标模型训练指标,基于目标模型训练指标和指标比例映射关系确定目标标签修正比例;特异性领域数据特征数量确定子单元,适于基于目标标签修正比例和领域数据特征集的领域数据特征数量,确定特异性领域数据特征数量;特异性领域数据特征确定子单元,适于基于特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
可选地,训练模块14包括:第一模型训练单元,适于基于目标领域模型训练数据集对生成式处理模型进行第一模型训练处理,得到第一生成式处理模型;检测单元,适于检测第一生成式处理模型是否满足模型结束训练条件;第一判断单元,适于若第一生成式处理模型不满足模型结束训练条件,则确定领域数据特征集中新增特异性领域数据特征,对新增特异性领域数据特征携带的特异性模型处理标签进行标签修正得到标签修正后的新增特异性领域数据特征,基于新增特异性领域数据特征对目标领域模型训练数据集进行更新得到更新后的目标领域模型训练数据集,基于目标领域模型训练数据集对第一生成式处理模型进行第二模型训练处理,得到第二生成式处理模型,将第二生成式处理模型作为第一生成式处理模型并行执行检测第一生成式处理模型是否满足模型结束训练条件的步骤;第二判断单元,适于若第一生成式处理模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
可选地,第一判断单元包括:新增特异性领域数据特征数量确定子单元,适于确定第一生成式处理模型的当前模型训练指标,基于当前模型训练指标与目标模型训练指标的指标差值确定新增特异性领域数据特征数量;新增特异性领域数据特征确定子单元,适于基于新增特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中新增特异性领域数据特征。
可选地,确定模块12包括:提取单元,适于提取目标领域数据的领域综合特征和领域属性特征,提取目标领域数据的通用领域特征;领域数据特征确定单元,适于基于领域综合特征和领域属性特征确定领域垂类特征,基于领域数据特征和通用领域数据特征确定领域数据特征。
可选地,领域数据特征确定单元包括:权重确定子单元,适于确定领域综合特征对应的第一权重系数,确定领域属性特征对应的第二权重系数;领域垂类特征确定子单元,适于将领域综合特征、领域属性特征、第一权重系数、第二权重系数采用第一计算式得到领域垂类特征;第一计算式满足以下公式:
FL=α*fc+β*fb
其中,FL是领域垂类特征,α是第一权重系数,β是第二权重系数,fc是领域综合特征,fb是领域属性特征;权重系数确定子单元,适于确定通用领域数据特征对应的第三权重系数,第一权重系数、第二权重系数和第三权重系数的和为1;领域数据特征子单元,适于将领域垂类特征、通用领域数据特征、第三权重系数采用第二计算式得到领域数据特征,基于领域数据特征确定领域数据特征;第二计算式满足以下公式:
FD=FL+γ*fi
其中,FD是领域数据特征,fi是通用领域数据特征,γ是第三权重系数。
可选地,获取模块11包括:生成式处理模型获取单元,适于获取针对目标领域的生成式处理模型;识别单元,适于获取源数据和针对目标领域的目标领域标识符,基于目标领域标识符从源数据中识别目标领域源数据;目标领域实体信息确定单元,适于获取目标领域实体定义信息,基于目标领域实体定义信息从目标领域源数据中进行实体信息提取,得到目标领域实体信息;目标领域数据提取单元,适于获取目标领域实体信息对应的数据提取逻辑,基于数据提取逻辑对目标领域实体信息进行数据提取,得到目标领域数据。
可选地,模型训练数据生成装置1还包括:特征相似度确定模块,适于确定目标领域模型训练数据集中两两领域数据特征之间的特征相似度;参考特征对确定模块,适于确定特征相似度大于或等于相似度阈值的参考特征相似度以及参考特征相似度对应的参考特征对;数据融合模块,适于确定参考特征对对应的两个相似领域数据特征,对相似领域数据特征进行数据融合处理,以更新目标领域模型训练数据集。
需要说明的是,上述实施例提供的模型训练数据生成装置在执行模型训练数据生成方法时,仅以上述各功能模块的划分进行举例说明,实际应用中,可以根据需要而将上述功能分配由不同的功能模块完成,即将设备的内部结构划分成不同的功能模块,以完成以上描述的全部或者部分功能。另外,上述实施例提供的模型训练数据生成装置与模型训练数据生成方法实施例属于同一构思,其体现实现过程详见方法实施例,这里不再赘述。
本说明书还提供了一种计算机存储介质,所述计算机存储介质可以存储有多条指令,所述指令适于由处理器加载并执行如上述图1~图10所示实施例的所述模型训练数据生成方法,具体执行过程可以参见图1~图10所示实施例的具体说明,在此不进行赘述。
本说明书还提供了一种计算机程序产品,该计算机程序产品存储有至少一条指令,所述至少一条指令由所述处理器加载并执行如上述图1~图10所示实施例的所述模型训练数据生成方法,具体执行过程可以参见图1~图10所示实施例的具体说明,在此不进行赘述。
请参考图12,图12为本说明书实施例提供的电子设备的结构方框图。本说明书中的电子设备可以包括一个或多个如下部件:处理器110、存储器120、输入装置130、输出装置140和总线150。处理器110、存储器120、输入装置130和输出装置140之间可以通过总线150连接。
处理器110可以包括一个或者多个处理核心。处理器110利用各种接口和线路连接整个电子设备内的各个部分,通过运行或执行存储在存储器120内的指令、程序、代码集或指令集,以及调用存储在存储器120内的数据,执行电子设备的各种功能和处理数据。可选地,处理器110可以采用数字信号处理(digital signal processing,DSP)、现场可编程门阵列(field-programmable gate array,FPGA)、可编程逻辑阵列(programmable logic Array,PLA)中的至少一种硬件形式来实现。处理器110可集成中心处理器(central processing unit,CPU)、图像处理器(graphics processing unit,GPU)和调制解调器等中的一种或几种的组合。其中,CPU主要处理操作系统、用户界面和应用程序等;GPU用于负责显示内容的渲染和绘制;调制解调器用于处理无线通信。可以理解的是,上述调制解调器也可以不集成到处理器110中,单独通过一块通信芯片进行实现。
存储器120可以包括随机存储器(random Access Memory,RAM),也可以包括只读存储器(read-only memory,ROM)。可选地,该存储器120包括非瞬时性计算机可读介质(non-transitory computer-readable storage medium)。存储器120可用于存储指令、程序、代码、代码集或指令集。存储器120可包括存储程序区和存储数据区,其中,存储程序区可存储用于实现操作系统的指令、用于实现至少一个功能的指令(比如触控功能、声音播放功能、图像播放功能等)、用于实现下述各个方法实施例的指令等,该操作系统可以是安卓(Android)系统,包括基于Android系统深度开发的系统、苹果公司开发的IOS系统,包括基于IOS系统深度开发的系统或其它系统。存储数据区还可以存储电子设备在使用中所创建的数据比如电话本、音视频数据、聊天记录数据,等。
参见图13所示,图13为本说明书实施例提供的一种操作系统和用户空间的结构示意图,存储器120可分为操作系统空间和用户空间,操作系统即运行于操作系统空间,原生及第三方应用程序即运行于用户空间。为了保证不同第三方应用程序均能够达到较好的运行效果,操作系统针对不同第三方应用程序为其分配相应的系统资源。然而,同一第三方应用程序中不同应用场景对系统资源的需求也存在差异,比如,在本地资源加载场景下,第三方应用程序对磁盘读取速度的要求较高;在动画渲染场景下,第三方应用程序则对GPU性能的要求较高。而操作系统与第三方应用程序之间相互独立,操作系统往往不能及时感知第三方应用程序当前的应用场景,导致操作系统无法根据第三方应用程序的具体应用场景进行针对性的系统资源适配。
为了使操作系统能够区分第三方应用程序的具体应用场景,需要打通第三方应用程序与操作系统之间的数据通信,使得操作系统能够随时获取第三方应用程序当前的场景信息,进而基于当前场景进行针对性的系统资源适配。
参见图14所示,图14为本说明书实施例提供的一种图13中安卓操作系统的架构图。以操作系统为Android系统为例,存储器120中存储的程序和数据如图14所示,存储器120中可存储有Linux内核层320、系统运行时库层340、应用框架层360和应用层380,其中,Linux内核层320、系统运行库层340和应用框架层360属于操作系统空间,应用层380属于用户空间。Linux内核层320为电子设备的各种硬件提供了底层的驱动,如显示驱动、音频驱动、摄像头驱动、蓝牙驱动、Wi-Fi驱动、电源管理等。系统运行库层340通过一些C/C++库来为Android系统提供了主要的特性支持。如SQLite库提供了数据库的支持,OpenGL/ES库提供了3D绘图的支持,Webkit库提供了浏览器内核的支持等。在系统运行时库层340中还提供有安卓运行时库(Android runtime),它主要提供了一些核心库,能够允许开发者使用Java语言来编写Android应用。应用框架层360提供了构建应用程序时可能用到的各种API,开发者也可以通过使用这些API来构建自己的应用程序,比如活动管理、窗口管理、视图管理、通知管理、内容提供者、包管理、通话管理、资源管理、定位管理。应用层380中运行有至少一个应用程序,这些应用程序可以是操作系统自带的原生应用程序,比如联系人程序、短信程序、时钟程序、相机应用等;也可以是第三方开发者所开发的第三方应用程序,比如游戏类应用程序、即时通信程序、相片美化程序等。
参见图15所示,图15为本说明书实施例提供的一种图13中IOS操作系统的架构图。以操作系统为IOS系统为例,存储器120中存储的程序和数据如图14所示,IOS系统包括:核心操作系统层420(Core OS layer)、核心服务层440(Core Services layer)、媒体层460(Media layer)、可触摸层480(Cocoa Touch Layer)。核心操作系统层420包括了操作系统内核、驱动程序以及底层程序框架,这些底层程序框架提供更接近硬件的功能,以供位于核心服务层440的程序框架所使用。核心服务层440提供给应用程序所需要的系统服务和/或程序框架,比如基础(Foundation)框架、账户框架、广告框架、数据存储框架、网络连接框架、地理位置框架、运动框架等等。媒体层460为应用程序提供有关视听方面的接口,如图形图像相关的接口、音频技术相关的接口、视频技术相关的接口、音视频传输技术的无线播放(AirPlay)接口等。可触摸层480为应用程序开发提供了各种常用的界面相关的框架,可触摸层480负责用户在电子设备上的触摸交互操作。比如本地通知服务、远程推送服务、广告框架、游戏工具框架、消息用户界面接口(User Interface,UI)框架、用户界面UIKit框架、地图框架等等。
在图15所示出的框架中,与大部分应用程序有关的框架包括但不限于:核心服务层440中的基础框架和可触摸层480中的UIKit框架。基础框架提供许多基本的对象类和数据类型,为所有应用程序提供最基本的系统服务,和UI无关。而UIKit框架提供的类是基础的UI类库,用于创建基于触摸的用户界面,iOS应用程序可以基于UIKit框架来提供UI,所以它提供了应用程序的基础架构,用于构建用户界面,绘图、处理和用户交互事件,响应手势等等。
其中,在IOS系统中实现第三方应用程序与操作系统数据通信的方式以及原理可参考Android系统,本说明书在此不再赘述。
其中,输入装置130用于接收输入的指令或数据,输入装置130包括但不限于键盘、鼠标、摄像头、麦克风或触控设备。输出装置140用于输出指令或数据,输出装置140包括但不限于显示设备和扬声器等。在一个示例中,输入装置130和输出装置140可以合设,输入装置130和输出装置140为触摸显示屏,该触摸显示屏用于接收用户使用手指、触摸笔等任何适合的物体在其上或附近的触摸操作,以及显示各个应用程序的用户界面。触摸显示屏通常设置在电子设备的前面板。触摸显示屏可被设计成为全面屏、曲面屏或异型屏。触摸显示屏还可被设计成为全面屏与曲面屏的结合,异型屏与曲面屏的结合,本说明书对此不加以限定。
除此之外,本领域技术人员可以理解,上述附图所示出的电子设备的结构并不构成对电子设备的限定,电子设备可以包括比图示更多或更少的部件,或者组合某些部件,或者不同的部件布置。比如,电子设备中还包括射频电路、输入单元、传感器、音频电路、无线保真(wireless fidelity,WiFi)模块、电源、蓝牙模块等部件,在此不再赘述。
在本说明书中,各步骤的执行主体可以是上文介绍的电子设备。可选地,各步骤的执行主体为电子设备的操作系统。操作系统可以是安卓系统,也可以是IOS系统,或者其它操作系统,本说明书对此不作限定。
本说明书的电子设备,其上还可以安装有显示设备,显示设备可以是各种能实现显示功能的设备,例如:阴极射线管显示器(cathode ray tubedisplay,简称CR)、发光二极管显示器(light-emitting diode display,简称LED)、电子墨水屏、液晶显示屏(liquid crystal display,简称LCD)、等离子显示面板(plasma display panel,简称PDP)等。用户可以利用电子设备101上的显示设备,来查看显示的文字、图像、视频等信息。所述电子设备可以是智能手机、平板电脑、游戏设备、AR(Augmented Reality,增强现实)设备、汽车、数据存储装置、音频播放装置、视频播放装置、笔记本、桌面计算设备、可穿戴设备诸如电子手表、电子眼镜、电子头盔、电子手链、电子项链、电子衣物等设备。
在图12所示的电子设备中,其中电子设备可以是一种终端,处理器110可以用于调用存储器120中存储的模型训练数据生成程序,并具体执行以下操作:获取生成式处理模型和目标领域数据;确定目标领域数据对应的领域数据特征以及领域数据特征的模型处理标签;基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;采用目标领域模型训练数据集对生成式处理模型进行模型训练。
可选地,处理器110执行基于领域数据特征对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集时,具体执行:基于领域数据特征确定特异性领域数据特征,对特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签;基于修正模型处理标签对领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集。
可选地,处理器110执行基于领域数据特征确定特异性领域数据特征时,具体执行:对领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于领域数据特征集确定特异性领域数据特征。
可选地,处理器110执行对领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于领域数据特征集确定特异性领域数据特征时,具体执行:计算两两领域数据特征之间的特征相似度,基于特征相似度进行聚类,得到至少两类领域数据特征集;基于领域数据特征集的所有特征相似度,确定特征集相似度分布信息;基于特征集相似度分布信息从领域数据特征集中确定特异性领域数据特征,基于特异性领域数据特征确定特异性领域数据特征。
可选地,处理器110执行基于特征集相似度分布信息从领域数据特征集中确定特异性领域数据特征时,具体执行:基于特征集相似度分布信息,从领域数据特征集中查询特征相似度小于或等于特征相似度阈值的目标特征相似度,确定目标特征相似度对应的领域数据特征对,基于领域数据特征对确定领域数据特征对集合;确定领域数据特征对集合中各参考领域数据特征的特征重复频次;基于各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
可选地,处理器110执行基于各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征时,具体执行:获取模型训练指标与标签修正比例的指标比例映射关系,确定针对生成式处理模型的目标模型训练指标,基于目标模型训练指标和指标比例映射关系确定目标标签修正比例;基于目标标签修正比例和领域数据特征集的领域数据特征数量,确定特异性领域数据特征数量;基于特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中的特异性领域数据特征。
可选地,处理器110执行采用目标领域模型训练数据集对生成式处理模型进行模型训练时,具体执行:基于目标领域模型训练数据集对生成式处理模型进行第一模型训练处理,得到第一生成式处理模型;检测第一生成式处理模型是否满足模型结束训练条件;若第一生成式处理模型不满足模型结束训练条件,则确定领域数据特征集中新增特异性领域数据特征,对新增特异性领域数据特征携带的特异性模型处理标签进行标签修正得到标签修正后的新增特异性领域数据特征,基于新增特异性领域数据特征对目标领域模型训练数据集进行更新得到更新后的目标领域模型训练数据集,基于目标领域模型训练数据集对第一生成式处理模型进行第二模型训练处理,得到第二生成式处理模型,将第二生成式处理模型作为第一生成式处理模型并行执行检测第一生成式处理模型是否满足模型结束训练条件的步骤;若第一生成式处理模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
可选地,处理器110执行确定领域数据特征集中新增特异性领域数据特征时,具体执行:确定第一生成式处理模型的当前模型训练指标,基于当前模型训练指标与目标模型训练指标的指标差值确定新增特异性领域数据特征数量;基于新增特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定领域数据特征集中新增特异性领域数据特征。
可选地,处理器110执行确定目标领域数据对应的领域数据特征时,具体执行:提取目标领域数据的领域综合特征和领域属性特征,提取目标领域数据的通用领域特征;基于领域综合特征和领域属性特征确定领域垂类特征,基于领域数据特征和通用领域数据特征确定领域数据特征。
可选地,处理器110执行基于领域综合特征和领域属性特征确定领域垂类特征,基于领域数据特征和通用领域数据特征确定领域数据特征时,具体执行:确定领域综合特征对应的第一权重系数,确定领域属性特征对应的第二权重系数;将领域综合特征、领域属性特征、第一权重系数、第二权重系数采用第一计算式得到领域垂类特征;第一计算式满足以下公式:
FL=α*fc+β*fb
其中,FL是领域垂类特征,α是第一权重系数,β是第二权重系数,fc是领域综合特征,fb是领域属性特征;确定通用领域数据特征对应的第三权重系数,第一权重系数、第二权重系数和第三权重系数的和为1;将领域垂类特征、通用领域数据特征、第三权重系数采用第二计算式得到领域数据特征,基于领域数据特征确定领域数据特征;第二计算式满足以下公式:
FD=FL+γ*fi
其中,FD是领域数据特征,fi是通用领域数据特征,γ是第三权重系数。
可选地,处理器110执行获取生成式处理模型和目标领域数据时,具体执行:获取针对目标领域的生成式处理模型;获取源数据和针对目标领域的目标领域标识符,基于目标领域标识符从源数据中识别目标领域源数据;获取目标领域实体定义信息,基于目标领域实体定义信息从目标领域源数据中进行实体信息提取,得到目标领域实体信息;获取目标领域实体信息对应的数据提取逻辑,基于数据提取逻辑对目标领域实体信息进行数据提取,得到目标领域数据。
可选地,处理器110还适于执行:确定目标领域模型训练数据集中两两领域数据特征之间的特征相似度;确定特征相似度大于或等于相似度阈值的参考特征相似度以及参考特征相似度对应的参考特征对;确定参考特征对对应的两个相似领域数据特征,对相似领域数据特征进行数据融合处理,以更新目标领域模型训练数据集。
本领域普通技术人员可以理解实现上述实施例方法中的全部或部分流程,是可以通过计算机程序来指令相关的硬件来完成,所述的程序可存储于一计算机可读取存储介质中,该程序在执行时,可包括如上述各方法的实施例的流程。其中,所述的存储介质可为磁碟、光盘、只读存储记忆体或随机存储记忆体等。
需要说明的是,本说明书实施例所涉及的信息(包括但不限于用户设备信息、用户个人信息等)、数据(包括但不限于用于分析的数据、存储的数据、展示的数据等)以及信号,均为经用户授权或者经过各方充分授权的,且相关数据的收集、使用和处理需要遵守相关国家和地区的相关法律法规和标准。例如,本说明书中涉及的对象特征、交互行为特征以及用户信息等都是在充分授权的情况下获取的。
以上所揭露的仅为本说明书较佳实施例而已,当然不能以此来限定本说明书之权利范围,因此依本说明书权利要求所作的等同变化,仍属本说明书所涵盖的范围。

Claims (16)

  1. 一种模型训练数据生成方法,所述方法包括:
    获取生成式处理模型和目标领域数据;
    确定所述目标领域数据对应的领域数据特征以及所述领域数据特征的模型处理标签;
    基于所述领域数据特征对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;
    采用所述目标领域模型训练数据集对所述生成式处理模型进行模型训练。
  2. 根据权利要求1所述的方法,所述基于所述领域数据特征对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集,包括:
    基于所述领域数据特征确定特异性领域数据特征,对所述特异性领域数据特征对应的特异性模型处理标签进行标签修正处理,得到修正模型处理标签;
    基于所述修正模型处理标签对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集。
  3. 根据权利要求2所述的方法,所述基于所述领域数据特征确定特异性领域数据特征,包括:
    对所述领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于所述领域数据特征集确定特异性领域数据特征。
  4. 根据权利要求3所述的方法,所述对所述领域数据特征进行聚类处理,得到至少两类领域数据特征集,基于所述领域数据特征集确定特异性领域数据特征,包括:
    计算两两所述领域数据特征之间的特征相似度,基于所述特征相似度进行聚类,得到至少两类领域数据特征集;
    基于所述领域数据特征集的所有所述特征相似度,确定特征集相似度分布信息;
    基于所述特征集相似度分布信息从所述领域数据特征集中确定特异性领域数据特征,基于所述特异性领域数据特征确定特异性领域数据特征。
  5. 根据权利要求4所述的方法,所述基于所述特征集相似度分布信息从所述领域数据特征集中确定特异性领域数据特征,包括:
    基于所述特征集相似度分布信息,从所述领域数据特征集中查询特征相似度小于或等于特征相似度阈值的目标特征相似度,确定所述目标特征相似度对应的领域数据特征对,基于所述领域数据特征对确定领域数据特征对集合;
    确定所述领域数据特征对集合中各参考领域数据特征的特征重复频次;
    基于各参考领域数据特征的特征重复频次,确定所述领域数据特征集中的特异性领域数据特征。
  6. 根据权利要求5所述的方法,所述基于各参考领域数据特征的特征重复频次,确定所述领域数据特征集中的特异性领域数据特征,包括:
    获取模型训练指标与标签修正比例的指标比例映射关系,确定针对所述生成式处理模型的目标模型训练指标,基于所述目标模型训练指标和所述指标比例映射关系确定目标标签修正比例;
    基于所述目标标签修正比例和所述领域数据特征集的领域数据特征数量,确定特异性领域数据特征数量;
    基于所述特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定所述领域数据特征集中的特异性领域数据特征。
  7. 根据权利要求3至6中任意一项所述的方法,所述采用所述目标领域模型训练数据集对所述生成式处理模型进行模型训练,包括:
    基于所述目标领域模型训练数据集对所述生成式处理模型进行第一模型训练处理,得到第一生成式处理模型;
    检测所述第一生成式处理模型是否满足模型结束训练条件;
    若所述第一生成式处理模型不满足模型结束训练条件,则确定所述领域数据特征集中新增特异性领域数据特征,对所述新增特异性领域数据特征携带的特异性模型处理标签进行标签修正得到标签修正后的所述新增特异性领域数据特征,基于所述新增特异性领域数据特征对所述目标领域模型训练数据集进行更新得到更新后的所述目标领域模型训练数据集,基于所述目标领域模型训练数据集对所述第一生成式处理模型进行第二模型训练处理,得到第二生成式处理模型,将所述第二生成式处理模型作为所述第一生成式处理模型并行执行所述检测所述第一生成式处理模型是否满足模型结束训练条件的步骤;
    若所述第一生成式处理模型满足模型结束训练条件,则确定结束模型训练得到目标生成式处理模型。
  8. 根据权利要求7所述的方法,所述确定所述领域数据特征集中新增特异性领域数据特征,包括:
    确定所述第一生成式处理模型的当前模型训练指标,基于所述当前模型训练指标与所述目标模型训练指标的指标差值确定新增特异性领域数据特征数量;
    基于所述新增特异性领域数据特征数量和各参考领域数据特征的特征重复频次,确定所述领域数据特征集中新增特异性领域数据特征。
  9. 根据权利要求1所述的方法,所述确定所述目标领域数据对应的领域数据特征,包括:
    提取所述目标领域数据的领域综合特征和领域属性特征,提取所述目标领域数据的通用领域特征;
    基于所述领域综合特征和所述领域属性特征确定领域垂类特征,基于所述领域数据特征和所述通用领域数据特征确定领域数据特征。
  10. 根据权利要求9所述的方法,所述基于所述领域综合特征和所述领域属性特征确定领域垂类特征,基于所述领域数据特征和所述通用领域数据特征确定领域数据特征,包括:
    确定所述领域综合特征对应的第一权重系数,确定所述领域属性特征对应的第二权重系数;
    将所述领域综合特征、所述领域属性特征、所述第一权重系数、所述第二权重系数采用第一计算式得到领域垂类特征;
    所述第一计算式满足以下公式:
    FL=α*fc+β*fb
    其中,FL是所述领域垂类特征,α是所述第一权重系数,β是所述第二权重系数,fc是所述领域综合特征,fb是所述领域属性特征;
    确定所述通用领域数据特征对应的第三权重系数,所述第一权重系数、所述第二权重系数和所述第三权重系数的和为1;
    将所述领域垂类特征、所述通用领域数据特征、所述第三权重系数采用第二计算式得到领域数据特征,基于所述领域数据特征确定领域数据特征;
    所述第二计算式满足以下公式:
    FD=FL+γ*fi
    其中,FD是所述领域数据特征,fi是所述通用领域数据特征,γ是所述第三权重系数。
  11. 根据权利要求1所述的方法,所述获取生成式处理模型和目标领域数据,包括:
    获取针对目标领域的生成式处理模型;
    获取源数据和针对目标领域的目标领域标识符,基于所述目标领域标识符从所述源数据中识别目标领域源数据;
    获取目标领域实体定义信息,基于所述目标领域实体定义信息从所述目标领域源数据中进行实体信息提取,得到目标领域实体信息;
    获取所述目标领域实体信息对应的数据提取逻辑,基于所述数据提取逻辑对所述目标领域实体信息进行数据提取,得到目标领域数据。
  12. 根据权利要求1所述的方法,所述得到目标领域模型训练数据集之后,所述方法还包括:
    确定所述目标领域模型训练数据集中两两所述领域数据特征之间的特征相似度;
    确定特征相似度大于或等于相似度阈值的参考特征相似度以及所述参考特征相似度对应的参考特征对;
    确定所述参考特征对对应的两个相似领域数据特征,对所述相似领域数据特征进行数据融合处理,以更新所述目标领域模型训练数据集。
  13. 一种模型训练数据生成装置,所述装置包括:
    获取模块,适于获取生成式处理模型和目标领域数据;
    确定模块,适于确定所述目标领域数据对应的领域数据特征以及所述领域数据特征的模型处理标签;
    更新模块,适于基于所述领域数据特征对所述领域数据特征的模型处理标签进行标签更新处理,得到目标领域模型训练数据集;
    训练模块,适于采用所述目标领域模型训练数据集对所述生成式处理模型进行模型训练。
  14. 一种计算机存储介质,所述计算机存储介质存储有多条指令,所述指令适于由处理器加载并执行如权利要求1~12任意一项的方法步骤。
  15. 一种计算机程序产品,该计算机程序产品存储有至少一条指令,所述至少一条指令由处理器加载并执行如权利要求1~12任意一项的方法步骤。
  16. 一种电子设备,包括:处理器和存储器;其中,所述存储器存储有计算机程序,所述计算机程序适于由所述处理器加载并执行如权利要求1~12任意一项的方法步骤。
PCT/CN2025/085413 2024-03-27 2025-03-27 模型训练数据生成 Pending WO2025201470A1 (zh)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
CN202410364205.1A CN118070923A (zh) 2024-03-27 2024-03-27 模型训练数据生成方法、装置、存储介质及电子设备
CN202410364205.1 2024-03-27

Publications (1)

Publication Number Publication Date
WO2025201470A1 true WO2025201470A1 (zh) 2025-10-02

Family

ID=91111328

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/CN2025/085413 Pending WO2025201470A1 (zh) 2024-03-27 2025-03-27 模型训练数据生成

Country Status (2)

Country Link
CN (1) CN118070923A (zh)
WO (1) WO2025201470A1 (zh)

Families Citing this family (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN118070923A (zh) * 2024-03-27 2024-05-24 支付宝(杭州)信息技术有限公司 模型训练数据生成方法、装置、存储介质及电子设备

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113378895A (zh) * 2021-05-24 2021-09-10 成都欧珀通信科技有限公司 一种分类模型生成方法、装置、存储介质及电子设备
CN113705554A (zh) * 2021-08-13 2021-11-26 北京百度网讯科技有限公司 图像识别模型的训练方法、装置、设备及存储介质
CN117172328A (zh) * 2023-08-22 2023-12-05 火离科技(上海)有限公司 一种垂类模型训练方法、装置、电子设备及可读存储介质
CN118070923A (zh) * 2024-03-27 2024-05-24 支付宝(杭州)信息技术有限公司 模型训练数据生成方法、装置、存储介质及电子设备

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113378895A (zh) * 2021-05-24 2021-09-10 成都欧珀通信科技有限公司 一种分类模型生成方法、装置、存储介质及电子设备
CN113705554A (zh) * 2021-08-13 2021-11-26 北京百度网讯科技有限公司 图像识别模型的训练方法、装置、设备及存储介质
CN117172328A (zh) * 2023-08-22 2023-12-05 火离科技(上海)有限公司 一种垂类模型训练方法、装置、电子设备及可读存储介质
CN118070923A (zh) * 2024-03-27 2024-05-24 支付宝(杭州)信息技术有限公司 模型训练数据生成方法、装置、存储介质及电子设备

Also Published As

Publication number Publication date
CN118070923A (zh) 2024-05-24

Similar Documents

Publication Publication Date Title
CN111382228B (zh) 用于输出信息的方法和装置
WO2024099457A1 (zh) 一种信息推荐方法、装置、存储介质及电子设备
CN111523021A (zh) 一种信息处理系统及其执行的方法
CN110457325B (zh) 用于输出信息的方法和装置
CN110275962B (zh) 用于输出信息的方法和装置
US20240152933A1 (en) Automatic mapping of a question or compliance controls associated with a compliance standard to compliance controls associated with another compliance standard
EP4134885A2 (en) Method and apparatus of recommending data, electronic device, and medium
CN115049068B (zh) 一种模型处理方法、装置、存储介质及电子设备
WO2025171816A1 (zh) 文本意图识别模型训练、文本意图识别
WO2025201472A1 (zh) 事实评估模型训练、事实评估
CN112631588A (zh) 文件生成方法、装置、电子设备和计算机可读介质
WO2025201470A1 (zh) 模型训练数据生成
CN112582073B (zh) 医疗信息获取方法、装置、电子设备和介质
CN116071527A (zh) 一种对象处理方法、装置、存储介质及电子设备
CN117056507B (zh) 长文本分析方法、长文本分析模型训练方法及相关设备
CN109857838B (zh) 用于生成信息的方法和装置
CN116522996A (zh) 一种推荐模型的训练方法、推荐方法以及相关装置
CN113807056B (zh) 一种文档名称序号纠错方法、装置和设备
CN115827833A (zh) 一种对话结构处理方法、装置、存储介质及电子设备
CN116228391A (zh) 一种风险识别方法、装置、存储介质及电子设备
CN111784377A (zh) 用于生成信息的方法和装置
CN116881465A (zh) 一种企业关系图谱生成方法、装置、存储介质及电子设备
CN114926234A (zh) 物品信息推送方法、装置、电子设备和计算机可读介质
CN115952403A (zh) 评估对象性能的方法、装置、电子设备及存储介质
CN116644331B (zh) 一种人群分类方法、装置、存储介质及电子设备

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 25775171

Country of ref document: EP

Kind code of ref document: A1