WO2020121115A1 - コンテンツの分類方法および分類モデルの生成方法 - Google Patents

コンテンツの分類方法および分類モデルの生成方法 Download PDF

Info

Publication number
WO2020121115A1
WO2020121115A1 PCT/IB2019/060377 IB2019060377W WO2020121115A1 WO 2020121115 A1 WO2020121115 A1 WO 2020121115A1 IB 2019060377 W IB2019060377 W IB 2019060377W WO 2020121115 A1 WO2020121115 A1 WO 2020121115A1
Authority
WO
WIPO (PCT)
Prior art keywords
content
classification
learning
classification model
contents
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/IB2019/060377
Other languages
English (en)
French (fr)
Inventor
桃純平
福留貴浩
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Semiconductor Energy Laboratory Co Ltd
Original Assignee
Semiconductor Energy Laboratory Co Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Semiconductor Energy Laboratory Co Ltd filed Critical Semiconductor Energy Laboratory Co Ltd
Priority to US17/311,730 priority Critical patent/US20220027799A1/en
Priority to DE112019006203.4T priority patent/DE112019006203T5/de
Priority to KR1020217016573A priority patent/KR20210100613A/ko
Priority to CN201980078452.2A priority patent/CN113168421A/zh
Priority to JP2020559054A priority patent/JP7730641B2/ja
Publication of WO2020121115A1 publication Critical patent/WO2020121115A1/ja
Anticipated expiration legal-status Critical
Priority to JP2024227415A priority patent/JP7734819B2/ja
Priority to JP2025140625A priority patent/JP2025161942A/ja
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/906Clustering; Classification
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/20Ensemble learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/21Design or setup of recognition systems or techniques; Extraction of features in feature space; Blind source separation
    • G06F18/214Generating training patterns; Bootstrap methods, e.g. bagging or boosting
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques
    • G06F18/241Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches
    • G06F18/2415Classification techniques relating to the classification model, e.g. parametric or non-parametric approaches based on parametric or probabilistic models, e.g. based on likelihood ratio or false acceptance rate versus a false rejection rate
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F18/00Pattern recognition
    • G06F18/20Analysing
    • G06F18/24Classification techniques
    • G06F18/243Classification techniques relating to the number of classes
    • G06F18/2431Multiple classes
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F3/00Input arrangements for transferring data to be processed into a form capable of being handled by the computer; Output arrangements for transferring data from processing unit to output unit, e.g. interface arrangements
    • G06F3/01Input arrangements or combined input and output arrangements for interaction between user and computer
    • G06F3/048Interaction techniques based on graphical user interfaces [GUI]
    • G06F3/0481Interaction techniques based on graphical user interfaces [GUI] based on specific properties of the displayed interaction object or a metaphor-based environment, e.g. interaction with desktop elements like windows or icons, or assisted by a cursor's changing behaviour or appearance
    • G06F3/0482Interaction with lists of selectable items, e.g. menus
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F9/00Arrangements for program control, e.g. control units
    • G06F9/06Arrangements for program control, e.g. control units using stored programs, i.e. using an internal store of processing equipment to receive or retain programs
    • G06F9/44Arrangements for executing specific programs
    • G06F9/451Execution arrangements for user interfaces
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • G06N20/10Machine learning using kernel methods, e.g. support vector machines [SVM]
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N5/00Computing arrangements using knowledge-based models
    • G06N5/01Dynamic search techniques; Heuristics; Dynamic trees; Branch-and-bound

Definitions

  • One aspect of the present invention relates to a content classification method using a computer device, a content classification system, a classification model generation method, and a graphical user interface.
  • one embodiment of the present invention relates to a computer device.
  • One aspect of the present invention relates to a method of classifying electronic contents (text data, image data, audio data, or moving image data) using a computer device.
  • one aspect of the present invention relates to a content classification system that efficiently classifies a collection of content using machine learning.
  • one aspect of the present invention relates to a content classification method, a content classification system, and a classification model generation method using a graphical user interface managed by a computer device by a program.
  • the user wants to easily classify and extract information about topics specified by the user from a collection of contents.
  • the result of classifying content varies depending on the knowledge or experience of individuals.
  • Patent Document 1 discloses a machine learning approach for determining a document highly relevant to a topic designated by a user.
  • the content is a patent.
  • Each patent is given a unique patent number. Therefore, hereinafter, the content may be described as a patent number. Note that in the content classification method handled in one aspect of the present invention, attention is paid to a plurality of management parameters given to patent numbers.
  • the content is not limited to patent documents.
  • the content can handle information such as text data, image data, audio data, or moving image data.
  • the meta information indicates not the content itself but data describing the attribute to which the content belongs or related information.
  • the patent number is associated with the claims, the abstract, the drawing, and the specification as the contents. Further, meta information (evaluation information, elapsed days, family information, etc.) is given to the patent number, and management is performed using the meta information. Patent numbers are classified according to their importance by using meta information. Although the accuracy and efficiency of classification depend on the content of the target document, it is easy to make a difference depending on the experience and skill of the user, and it is necessary to classify a large number of documents.
  • an object of one embodiment of the present invention is to provide a method for efficiently generating a classification model and classifying information using the classification model.
  • a program is stored in the storage device of the computer device.
  • the program can display various information on a display device included in the computer device through a graphical user interface (hereinafter, GUI).
  • GUI graphical user interface
  • the user can operate a computer, give information, respond to a database, instruct a machine learning, etc. on a computer device via a GUI.
  • the program can display the calculation processing result by machine learning, the content of learning content or unclassified content downloaded from the database on the display device via the GUI.
  • content when simply referred to as content, it includes learning content, unclassified content, or classified content.
  • the proposed content classification system uses machine learning to generate a content classification model, and uses the generated content classification model to classify unclassified content. For example, the content having a plurality of pieces of meta information is used as the learning content.
  • the learning content is further provided with a learning label to generate a feature vector from the learning content.
  • the meta information or the learning label can be treated as the feature amount of the learning content.
  • the learning contents are treated as teacher data.
  • the classification model can be acquired by performing machine learning based on the learning content.
  • the classification model obtained here classifies contents having a plurality of pieces of meta information.
  • the types of classification may be two or three or more depending on the purpose of the user.
  • the learning content can be downloaded from the learning content stored in the database.
  • the learning content stored in the storage device of the computer device can be used.
  • the learning content may be managed by including a learning label.
  • the classification model stored in the database may be downloaded.
  • a classification model stored in the storage device of the computer device may be used.
  • One aspect of the present invention includes a learning content and a content, wherein the learning content is provided with a first feature amount and a learning label, and the content is provided with a second feature amount. It A step of generating a plurality of first classification models by machine learning using a plurality of learning contents; a step of generating a second classification model using a plurality of first classification models; and a second classification model Is used to add determination information to a plurality of contents and display the contents in a graphical user interface.
  • One aspect of the present invention includes a learning content and a content, wherein the learning content is provided with a first feature amount and a learning label, and the content is provided with a second feature amount. It A step of generating a plurality of first classification models by machine learning using a plurality of learning contents, a step of calculating an average value from the outputs of the plurality of first classification models, and a step of using a plurality of average values A method of classifying contents, including a step of generating a second classification model, and a step of adding determination information to a plurality of contents using the second classification model and displaying the contents in a graphical user interface.
  • One aspect of the present invention includes a learning content and a content, wherein the learning content is provided with a first feature amount and a learning label, and the content is provided with a second feature amount. It A step of generating a plurality of first classification models by machine learning using a plurality of learning contents; a step of evaluating each of the plurality of first classification models according to a first evaluation criterion; Each of the classification models performs an evaluation based on the second evaluation criteria, an evaluation result based on the plurality of first evaluation criteria, and a step of generating a second classification model from the evaluation results based on the second evaluation criteria, And a step of adding determination information to a plurality of contents by using a second classification model and displaying the plurality of contents in a graphical user interface.
  • the first evaluation criterion is accuracy and the second evaluation criterion is sensitivity.
  • a content classification method including a step of generating a first classification model using arbitrary learning content is preferable.
  • the learning content is further provided with classification information, and the content having the same judgment information as the classification information is selected from the plurality of contents to which the classification label is given using the output of the second classification model. And displaying in a graphical user interface.
  • the learning content or the feature amount given to the content is preferably the content classification method which is a management parameter.
  • the determination information is preferably a content classification method including a classification label or score.
  • the graphical user interface includes a step of designating a specific numerical range in the score and displaying the corresponding content as a list.
  • One aspect of the present invention can provide a method of accurately classifying information.
  • one embodiment of the present invention can provide a user interface for accurately classifying information.
  • one embodiment of the present invention can provide a program for classifying information with high accuracy.
  • an interactive interface for generating a classification model using machine learning can be provided to a user, and the burden on the user such as preparation of teacher data and evaluation of learning results can be reduced. be able to.
  • the effects of one aspect of the present invention are not limited to the effects listed above.
  • the effects listed above do not prevent the existence of other effects.
  • the other effects are the effects which are not mentioned in this item, which will be described below.
  • the effects not mentioned in this item can be derived from the description such as the specification or the drawings by those skilled in the art, and can be appropriately extracted from these descriptions.
  • one embodiment of the present invention has at least one of the effects listed above and/or other effects. Therefore, one embodiment of the present invention may not have the effects listed above in some cases.
  • FIG. 1 is a flowchart illustrating a classification method.
  • FIG. 2 is a flowchart illustrating the classification method.
  • FIG. 3 is a diagram for explaining the connection between the classification system 100 and the network.
  • FIG. 4 is a block diagram illustrating the classification system.
  • 5A and 5B are diagrams illustrating a graphical user interface.
  • FIG. 6 is a diagram illustrating a method of generating a classification model.
  • FIG. 7 is a diagram illustrating a method of generating a classification model.
  • FIG. 8 is a diagram illustrating a method of generating a classification model.
  • FIG. 9 is a diagram illustrating a graphical user interface.
  • FIG. 10 is a diagram illustrating a graphical user interface.
  • the content classification method described in this embodiment is controlled by a program running on a computer device.
  • the program is stored in the memory or storage of the computer device. Alternatively, it is stored in a computer connected via a network (LAN (Local Area Network), WAN (Wide Area Network), the Internet, etc.) or a server computer having a database.
  • LAN Local Area Network
  • WAN Wide Area Network
  • the Internet etc.
  • server computer having a database.
  • the display device included in the computer device can display the data given to the program by the user and the result of calculation of the data by the arithmetic device included in the computer device.
  • the configuration of the device will be described in detail with reference to FIG.
  • the data displayed on the display device can be easily recognized by the user and the operability can be improved by, for example, following the listed display format. Therefore, a GUI is used as an interface for a user to easily interact with a program included in a computer device via a display device.
  • the user can use the content classification method of the program via the GUI.
  • the user can simplify the content classification operation using the GUI. Further, the user can easily visually judge the classification result of the content through the GUI. In addition, the user can easily operate the program through the GUI.
  • the content indicates information such as text data, image data, audio data, or moving image data.
  • the data processing unit has a data collection unit and a data generation unit.
  • the data collection unit acquires a file including a plurality of contents from the database via the GUI.
  • the data generation unit can generate the learning content by the user giving the learning label to the content via the GUI.
  • the learning content to which the learning label is attached may be acquired from the database.
  • the plurality of contents are files stored in the memory or storage of the computer device, or data stored in a database connected to a network, a computer, a data server, or the like.
  • a plurality of learning contents or a plurality of unclassified contents be listed and stored in the database.
  • a plurality of feature amounts and learning labels are given to the learning content and the unclassified content.
  • the learning label can be modified by the user via the GUI.
  • the learning label is attached to the learning content, the learning content to which the learning label is attached can be stored in the database.
  • -Learning contents can include verification contents without learning labels.
  • the verification content can be used to verify the classification model generated using the learning content.
  • the meta information includes evaluation information, number of days elapsed, number of families, family status, application type, life, number of pending in family, number of abandoned in family, cost, number of inventors, fields, or number of claims. and so on. That is, the meta information is a content management parameter.
  • the family includes a patent family, a patent family, and the like.
  • the learning processing unit has a step of generating a classification model using the learning content.
  • the learning processing unit has a classification model generation unit or a classification model evaluation unit.
  • the classification model generator can generate a classification model.
  • the classification model generation unit includes a step of generating a plurality of first classification models by machine learning using a plurality of learning contents, and generates a second classification model using the plurality of first classification models. Including steps.
  • the output value of the first classification model or the second classification model can be displayed on the GUI.
  • the user can give (including correction) the learning label of the first classification model to the output value. Alternatively, the user can add new learning content for the output value.
  • the classification model evaluation unit evaluates the classification model generated by the classification model generation unit using the verification content.
  • the classification model outputs the inferred result as determination information.
  • the GUI can display the evaluation content with the determination information added thereto.
  • the user can judge the output result of the classification model evaluation unit, correct the learning label if necessary, and update the classification model in the classification model generation unit.
  • the classification model can be updated by the classification model generation unit by adding learning contents.
  • the determination processing unit has a classification inference unit and a list generation unit.
  • the classification inference unit infers and classifies a plurality of unclassified contents using the first learning model and the second learning model generated by the classification model generation unit.
  • the classification model attaches the inference result to each content as determination information.
  • the list generation unit can generate a list in a format required by the user from the content provided with the determination information and display the list on the GUI. For example, when each content is managed for each application country, the application country can be used as the classification information. When the application country is used as the classification information, it is preferable that the generated classification models generate different classification models for each application country.
  • the classification information is not limited to the application country. For example, one of the meta information included in the content can be used as the classification information.
  • meta information is used as classification information.
  • the status of the patent family may be used as the classification information.
  • the patent number of the parent application, the patent number of the divisional application, and the like are given as meta information to the patent number.
  • a state in which a divisional application is possible from the parent application's patent number a state in which a divisional application is not possible from the parent application's patent number, a state in which the parent application's patent number is pending, a state in which the parent application's patent number is lost,
  • the division number can be further divided, the division number cannot be further divided, the division application cannot be divided, the division application patent number is pending, or the division application patent number is lost.
  • Different classification models can be generated depending on the state and the like, and the respective classification models can be used to make inferences.
  • the judgment processing unit can infer a plurality of unclassified contents by using the classification model.
  • the inferred result includes a step of adding the determination information to each content and displaying the determination information on the GUI.
  • the determination information includes at least a classification label and a score (probability).
  • the GUI includes a step of designating a specific numerical range of the score and displaying the corresponding content as a list.
  • the classification model generation unit generates a plurality of first classification models by machine learning using a plurality of learning contents; calculates an average value from outputs of the plurality of first classification models; Generating a second classification model using the average value of.
  • the output value of the first classification model or the second classification model can be displayed on the GUI.
  • the user can correct the learning label of the first classification model for the output value.
  • the user can add learning content to the output value.
  • the said average value means calculating using any one of an arithmetic mean calculation, a geometric mean calculation, or a harmonic mean calculation.
  • the second classification model is generated using the multiple average values.
  • the second classification model since the outputs of the first classification model are averaged, it is possible to reduce the influence of noise components such as outliers included in the learning content.
  • the classification model generation unit includes a step of generating a plurality of first classification models by machine learning using a plurality of learning contents, and a step of evaluating each of the plurality of first classification models according to a first evaluation criterion.
  • a step of each of the plurality of first classification models performing evaluation based on the second evaluation criterion, and a second classification model generated from the evaluation results based on the plurality of first evaluation criteria and the second evaluation criteria
  • the output value of the first classification model or the second classification model can be displayed on the GUI.
  • the user can correct the learning label of the first classification model for the output value. Alternatively, the user can add learning content to the output value.
  • the first evaluation criterion is the accuracy of the confusion matrix
  • the second evaluation criterion is the sensitivity of the confusion matrix.
  • the generated classification model can include the precision and recall of the plurality of first classification models. The classification accuracy of the second classification model generated using the plurality of first classification models is improved.
  • the second classification model may be generated by using the generated m first classification models (m represents a natural number).
  • the classification model generation unit described above has k learning contents (k represents a natural number)
  • the first classification model can generate k learning contents or less.
  • any learning content can include two different learning models.
  • the learning content to which k different numbers are given can generate the first classification model by using q (q represents a natural number) each in accordance with the order sorted by the numbers.
  • the program can display the content read from the database on the GUI.
  • the content preferably has listed meta-information.
  • the GUI displays the content according to the display format of the GUI.
  • the listed meta information added to the content is preferably managed in units called records. For example, each record is composed of an ID (Identification) associated with a number, content (image data, audio data, or moving image data), meta information, or the like.
  • machine learning is focused on meta information, and a classification model is generated by the machine learning.
  • the classification model analyzes the meta information and classifies the content that has been converted into a feature vector.
  • the classification model may be an algorithm such as K-means or DBSCAN (density-based spatial clustering of applications with noise).
  • the program can generate a classification model using machine learning using learning contents to which a plurality of meta information and learning labels are added.
  • a classification model an algorithm such as a decision tree, naive Bayes, KNN (k Nearest Neighbor), SVM (Support Vector Machines), perceptron, logistic regression, or a neural network can be used.
  • the program can switch the classification model according to the number of learning contents.
  • a decision tree naive Bayes, logistic regression may be used when the number of learning contents is small
  • SVM a random forest
  • a neural network may be used when the number of learning contents is a certain amount or more.
  • the classification model used in this embodiment uses a random forest, which is one of the decision tree algorithms.
  • random sampling or cross variation can be used as the selection method of the meta information, the selection method of the learning content, or the selection method of the first classification model.
  • q each can be selected according to the order in which the assigned numbers are sorted.
  • FIG. 1 is a flowchart illustrating a content classification method according to one aspect of the present embodiment.
  • the content classification method is controlled by a program running on a computer. Therefore, the content can be classified by the program having the data processing unit, the learning processing unit, or the determination processing unit.
  • the program can classify the content desired by the user via the GUI. That is, the contents processed by the respective processing units described above correspond to the steps of the program.
  • step S11 the user can instruct through the GUI to load the file containing the content.
  • the file is stored in the database of the data processing unit. Note that the file includes learning content, unclassified content, and the like.
  • a plurality of learning contents or a plurality of unclassified contents be listed and stored in the database.
  • the user can add or correct the learning label of the learning content displayed on the GUI.
  • the file can include verification content to which a learning label is not attached.
  • Step S12 is a learning processing unit that generates a classification model using the loaded file.
  • the generated classification model can evaluate the verification content and display the evaluation result on the GUI.
  • the user can instruct the evaluation result to correct the learning label, add learning content, and the like.
  • the user can update the meta information by predicting the change over time of the meta information that the learning content has. If the user updates the meta information, the classification model can include changes in the classification model over time. Therefore, the user can obtain a change in the classification of the content over time.
  • the classification model can classify a content group whose value is expected to increase or a content group whose value is expected to decrease.
  • Step S13 is a determination processing unit.
  • Pure content is inferred using the classification model generated in step S12.
  • the classification model can add determination information to unclassified content based on the inference result.
  • the determination processing unit can display the content provided with the determination information on the GUI in a format required by the user.
  • the determination information includes at least a classification label and a score.
  • the GUI can specify a specific numerical value range of the score and display the corresponding content.
  • step S11 includes the data collection unit of step S21 and the data generation unit of step S22.
  • the data collection unit in step S21 can load the file from the database.
  • the meta information, the content, and the like can be managed by different databases.
  • the meta information may differ in the company, organization, or user who handles the content. Therefore, the data collection unit has a function of collecting meta-information about contents from different databases.
  • each database can be installed in different buildings, different areas, or different countries.
  • the data generation unit can manage contents and meta information in units called records.
  • each record is composed of an ID associated with a number, content (image data, audio data, or moving image data), meta information, or the like.
  • the user can generate the learning content by adding the learning label to the content displayed on the GUI.
  • the learning processing unit in step S12 has a classification model generation unit in step S23, a classification model evaluation unit in step S24, and an output result determination process in step S25.
  • the classification model generation unit can generate a classification model of content.
  • the classification model generation unit can generate a plurality of first classification models by machine learning using a plurality of learning contents.
  • a second classification model can be generated using the plurality of first classification models.
  • the GUI can display the output value of the first classification model or the second classification model.
  • the user can add (including correction) the learning label of the first classification model to the output value.
  • the user can add new learning content for the output value. It is possible to predict the change with time of the meta information included in the learning content and update the meta information. The effect obtained by updating the meta information by the user can be referred to the description of step S12.
  • the classification model evaluation unit can evaluate the classification model generated by the classification model generation unit by using the verification content.
  • the classification model outputs the result of inferring the verification content as the determination information.
  • the GUI can display the evaluation content with the determination information added thereto.
  • step S25 the user can determine the output result of the classification model evaluation unit in step S24, and determine that the content classification model has been sufficiently learned.
  • the user gives a GUI model generation completion (OK) instruction.
  • the user can determine that the content classification model is not sufficiently learned (NG).
  • the user returns to step S23 and updates the classification model by changing the learning label, adding learning content, or updating meta information.
  • the determination processing unit in step S13 includes the classification inference unit in step S26 and the list creation unit in step S27.
  • the classification inference unit infers and classifies a plurality of unclassified contents using the first learning model and the second learning model generated by the classification model generation unit.
  • the classification inference unit is given the unclassified content generated by the data generation unit in step S22.
  • the classification model attaches the inference result to each content as determination information.
  • the list generation unit can list the contents to which the determination information is given in a format required by the user and display the list on the GUI.
  • classification information different from the meta information may be given to each content. For example, when classification information is given to the learning content, different classification models can be generated for each classification information. Alternatively, one of the meta information included in the content can be used as the classification information.
  • the judgment information includes at least a classification label and a score.
  • the GUI can specify a specific numerical range of the score, list corresponding contents, and display the list on the GUI.
  • FIG. 3 is a diagram for explaining the connection between the classification system 100 having the content classification method described above and the network (NetWork).
  • the classification system 100 is connected to the communication network LAN1.
  • a database DB1 or client computers CL1 to CLn (n is a natural number) is connected to the communication network LAN1.
  • the communication network LAN1 can be connected to the communication network LAN2 via the network.
  • the network can use the Internet, the communication network WAN, or satellite communication.
  • a database DB2, client computers CL11 to CL1n, etc. are connected to the communication network LAN2.
  • the classification system 100 uses a file including content stored in the database DB1, the database DB2, the client computers CL1 to CLn, or the client computers CL11 to CL1n to generate content, classify content, generate models, and unclassify. Content can be classified.
  • the user can give an instruction to the GUI from the program operating in the classification system 100.
  • a user can classify the unclassified content by generating the classification model described above using information of databases installed in different countries via the Internet. That is, the content or meta information may be stored in different databases or client computers.
  • the GUI can display the classification results classified by the classification system 100 stored in the storage device of the computer device of the database DB1, the database DB2, the client computers CL1 to CLn, or the client computers CL11 to CL1n.
  • FIG. 4 is a block diagram illustrating the classification system 100 described in FIG.
  • the classification system 100 includes a GUI (Graphical User Interface) 110, a calculation unit 120, and a storage unit 130.
  • the GUI 110 has an input unit 111 and an output unit 112.
  • the input unit 111 has a function of selecting a content load source and a function of inputting a learning label.
  • the output unit 112 has a function of displaying a content list loaded from a database or the like and a function of displaying determination information output by the classification model.
  • the meta information included in the displayed content can be modified by the user via the GUI.
  • the calculation unit 120 has a data processing unit 121, a learning processing unit 122, and a determination processing unit 123.
  • the data processing unit 121 has a data collection unit and a data generation unit.
  • the learning processing unit 122 has a classification model generation unit that creates a classification model and a classification model evaluation unit that classifies the classification model.
  • the output result of the classification model evaluation unit has a function of evaluation result determination processing in which the user makes a determination.
  • the determination processing unit 123 has a classification inference unit and an output list creation unit that lists the results classified by the classification inference unit.
  • the arithmetic unit 120 uses a microprocessor to perform arithmetic processing of a program stored in a storage unit included in the computer device. However, the program can be processed using a DSP (Digital signal Processor) or a GPU (Graphics Processing Unit).
  • DSP Digital signal Processor
  • GPU Graphics Processing Unit
  • the storage unit 130 temporarily stores a list of content and meta information generated by loading from a database or the like.
  • a DRAM dynamic random access memory
  • 1T transistor
  • 1C capacity
  • An OS transistor may be used as the transistor used in the memory cell of the DRAM.
  • the OS transistor is a transistor including a metal oxide in a semiconductor layer.
  • a memory device in which an OS transistor is used for a memory cell is called an “OS memory”.
  • OS memory a RAM having 1T1C type memory cells is referred to as "DOSRAM (Dynamic Oxide Semiconductor RAM)".
  • the off current of the OS transistor is very small. Therefore, in the DOSRAM, the frequency of refresh can be reduced, so that the power required for the refresh operation can be reduced.
  • the off-state current referred to here is a current flowing between the source and the drain when the transistor is off. In the case where the transistor is an n-channel type, for example, when the threshold voltage is about 0 V to 2 V, the current flowing between the source and the drain when the voltage between the gate and the source is a negative voltage is an off current. Can be called.
  • FIG. 5A is a diagram illustrating the configuration of the GUI 30.
  • the GUI 30 shows, as an example, a management screen that displays a list of p learning contents.
  • the learning content is managed in record units.
  • the record includes a number (No) 31, a content (ID) 32, meta information (Feature) 33 (meta information (F1) 33a to meta information (Fm) 33m) indicating a feature amount, and classification information (Case) 34 ( The classification information (C1) 34a to the classification information (Cq) 34q), the learning label (J-Label) 35, and the like are included.
  • the learning label 35 gives one of two values of “Yes” and “No”, but the learning label 35 is not limited to two values and may be three or more values. ..
  • FIG. 5B is a diagram illustrating the configuration of the GUI 30A.
  • the GUI 30A shows a management screen that infers n unclassified contents by the evaluation inference unit and displays a list of the inferred determination information.
  • the unclassified content has the number 31, the content 32, the meta information 33, and the classification information 34, like the learning content.
  • a classification label (A-Label) 36 and a score (Score) 37 are added to each record as determination information.
  • the GUI 30 and the GUI 30A can be managed on the same display screen.
  • FIG. 9 or FIG. 10 described later a display example of a GUI that can display the learning content and the determination information on the same management screen is shown.
  • FIG. 6 is a diagram illustrating a method of generating a classification model using a plurality of feature features associated with the learning content Sample described above by machine learning. Each of the characteristics Feature indicates any one of the meta information, and corresponds to a management parameter for managing the content. Note that in this embodiment, a method of generating a classification model will be described using the calculation unit F, the calculation unit S, the calculation unit V, the first classification model, and the second classification model.
  • Each of the learning content Sample(1) to the learning content Sample(k) is provided with j features Feature and a learning label Label.
  • the calculation unit F1 can generate a feature vector Vlabel1(1) in a format that can be processed by a computer from the learning content Sample(1).
  • the calculation unit Fk can generate the feature vector Vlabel1(k) in a computer processable form from the learning content Sample(k).
  • the feature vector Vlabel1(1) can be generated by the calculation unit F1 by giving different weighting factors to the respective features. Further, the feature vector Vlabel1(1) can be generated by using j or less feature features selected at random.
  • the calculation units S1 to Sm correspond to different first classification models.
  • the calculation unit S1 can generate the first classification model by using the feature vector Vlabel1(1) to the feature vector Vlabel1(k).
  • the number of feature vectors Vlabel1 given to the calculation unit S1 may be k or less.
  • the calculation unit Sm can generate the first classification model using different feature vectors Vlabel1(1) to Vlabel1(k). Therefore, the two different first classification models each include k or less feature vectors Vlabel1 and any one of the feature vectors Vlabel1 can include the same feature vector.
  • the k pieces of learning content Sample that are selected to generate the first classification model may be selected at random, or may be selected in the order sorted by the number assigned to the learning content. ..
  • the first classification model may include variation in the learning content.
  • the tendency along the number assigned based on one of the characteristics of the time series or the meta information can be included.
  • the first classification model can generate the feature vector Vlabel2 using the feature vector Vlabel1 generated from the learning content Sample(1) to the learning content Sample(k).
  • the second classification model is generated by the calculation unit V1.
  • the calculation unit V1 has a step of generating a second classification model using m feature vectors Vlabel2.
  • the second classification model can generate classification models having different features by using the feature vector Vlabel2(1) to the feature vector Vlabel2(m).
  • the second classification model can output the output value POUT using the feature vector Vlabel1 generated from the learning content Sample(1) to the learning content Sample(k).
  • the GUI can display the output value POUT.
  • the output value POUT includes a classification label that is determination information and a score. Therefore, the second classification model can classify contents. In addition, the second classification model can add determination information to each content.
  • the judgment result is obtained by giving unclassified contents to the learning content Sample of the classification model.
  • the learning label is not attached to the unclassified content.
  • FIG. 7 is a diagram illustrating a method of generating a classification model different from that of FIG.
  • FIG. 7 different points from FIG. 6 will be described, and in the configuration of the invention (or the configuration of the embodiment), the same reference numerals are commonly used in different drawings for the same portions or portions having similar functions. The repeated description is omitted.
  • the average value Av of the m feature vectors Vlabel2 is calculated, and the feature vector Vlabel_a is generated.
  • the second classification model can be generated by using the p feature vectors Vlabel_a.
  • the second classification model can generate classification models having different characteristics by calculating the average value Av of the m feature vectors Vlabel2.
  • the generated classification model can make the classification of content accurate.
  • FIG. 8 is a diagram illustrating a method of generating a classification model different from that of FIG. 7.
  • FIG. 8 different points from FIG. 7 will be described, and in the configuration of the invention (or the configuration of the embodiment), the same reference numeral is commonly used in different drawings for the same portion or a portion having a similar function. The repeated description is omitted.
  • the evaluation criterion for evaluating the m feature vectors Vlabel2 is given to the evaluation determination unit JG.
  • the evaluation determination unit JG1 is given a precision (Precision) as the first evaluation criterion and can evaluate each feature vector Vlabel2(1).
  • the evaluation determination unit JG1 can evaluate the respective feature vectors Vlabel2(1) given the sensitivity as a second evaluation criterion.
  • the evaluation determination unit JG1 outputs the evaluation result Vlabel_b(1).
  • the second classification model is generated by using the evaluation result Vlabel_b(1) to the evaluation result Vlabel_b(p).
  • the plurality of feature vectors Vlabel2 may be evaluated by different first evaluation criteria and second evaluation criteria, or may be evaluated by the same evaluation criteria.
  • an average value of the evaluation results Vlabel_b based on the first evaluation standard and the second evaluation standard can be calculated as in FIG. 7.
  • the second classification model can generate classification models having different features by using the evaluation results of the m feature vectors Vlabel2.
  • the generated classification model can make the classification of content accurate.
  • FIG. 9 is a diagram illustrating the GUI 50.
  • the GUI 50 is a display area for content (learning content, unclassified content, classified content), an icon 58a for selecting a download source of a file including the content, and text for displaying address information in which the selected file is stored. It has a box 58b and an icon (Learning Start) 59 for executing machine learning.
  • Each record has a number (No) 51, an ID (Index) 52, a feature amount (Feature) 53, classification information (Case) 54, a learning label (JL) 55, a classification label (AL) 56, and a score (Prob. ) 57 components.
  • the feature amount 53 can display the feature amount F(1) 53a to the feature amount F(j) 53j as detailed information.
  • j is a natural number.
  • the classification information 54 can display classification information C(1) 54a to classification information C(4) 54d as detailed information.
  • the classification information can have a type that can be represented by a natural number.
  • FIG. 9 shows an example in which the result of classifying the learning content and the unclassified content by the classification model is displayed on the GUI 50.
  • the record numbers No. 1 to No. 3 correspond to learning contents.
  • a learning label is given to the learning content, and classification information is given to the record numbers No1 to No3.
  • Record numbers No4 to No8 correspond to the classified contents.
  • a classification label 56 and a score 57 are given to the classified contents.
  • FIG. 9 the result of classifying the record numbers No4 to No7 using the classification model obtained by learning the record numbers No1 and No3 is displayed.
  • the result of classifying the record number No8 is displayed using the classification model obtained by learning the record number No2.
  • the number of records is displayed up to 8 due to the space, but the number of records can handle a plurality of types.
  • the classification label 56 or the score 57 preferably has a sorting function.
  • the GUI can select and display the determination result in which the classification label 56 is “Yes”. Further, the GUI can specify and display the numerical range of the score 57.
  • the GUI can classify and display contents having the same characteristics as the learning contents to which the teacher data is given.
  • the above classification model can add judgment information to unclassified contents.
  • the classification label 56 and the score 57 are displayed in the determination information.
  • the user gives “No” to the classification label 56 by using the sort function.
  • the score 57 is set to "0.8" to "1.0".
  • the GUI can select and display a record having the same characteristics as the learning content whose patent number has been abandoned.
  • FIG. 10 is a diagram illustrating a GUI 50A different from FIG. FIG. 10 is an efficient GUI display example when a large number of records are handled. Note that in FIG. 10, points different from FIG. 9 are described, and in the configuration of the invention (or the configuration of the embodiment), the same portions or portions having similar functions are denoted by the same reference numerals in different drawings. , And the repeated description thereof will be omitted.
  • FIG. 10 is different from FIG. 9 in that records can be classified and displayed for any selected classification information.
  • the display can be switched according to the type of the classification information C(1) to the classification information C(4).
  • the user sees the plurality of feature amounts 53 given to the record and the judgment information of the classification model, and if sufficient classification accuracy is obtained, the updating of the classification model ends.
  • the determination information attached to the record and determines that the classification accuracy is not sufficient the learning label is attached to the record to which the user-specified label is not attached, and the icon 59 is pressed.
  • the classification model can be updated.
  • the feature amount 53 may be updated by predicting the change over time in the meta information included in the learning content.
  • the classification model can include changes in the classification model over time. Therefore, the user can obtain a change in the classification of the content over time.
  • the classification model can classify the content group whose value is expected to increase or the content group whose value is expected to decrease.
  • the feature quantity 53, the classification information 54a to 54d, the learning label 55, the classification label 56, or the numerical values or label information included in the score 57 can be displayed in a different order.
  • the selected numerical values and label information can be sorted and displayed in the required order by using the filter function. Thereby, the user can efficiently evaluate the determination result of the classification model.
  • the content classification method described with reference to FIGS. 1 to 10 can provide a method of classifying information having a high probability.
  • GUI is suitable for classifying highly probable information.
  • the program can update the classification model by giving new teaching data (label for learning) to the classification model.
  • the program can classify information with high probability by updating the classification model.
  • the generated classification model can be saved in the main body of the electronic device or in the external memory, and can be called and used when classifying a new file. Furthermore, the classification model can be updated according to the method described above while adding new teacher data.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Software Systems (AREA)
  • General Physics & Mathematics (AREA)
  • Artificial Intelligence (AREA)
  • Evolutionary Computation (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Databases & Information Systems (AREA)
  • Medical Informatics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Evolutionary Biology (AREA)
  • Human Computer Interaction (AREA)
  • Health & Medical Sciences (AREA)
  • Molecular Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Biophysics (AREA)
  • Biomedical Technology (AREA)
  • Probability & Statistics with Applications (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

コンテンツを分類する分類モデルを提供する。 学習用コンテンツと、コンテンツと、を有し、学習用コンテンツには、第1の特徴量および学習用 ラベルが付与され、コンテンツには、第2の特徴量が付与される。複数の学習用コンテンツを用い て複数の第1の分類モデルを機械学習によって生成するステップと、複数の第1の分類モデルを用 いて第2の分類モデルを生成するステップと、第2の分類モデルを用いて複数のコンテンツに判定 情報を付与しGUIに表示するステップと、を含むコンテンツの分類方法である。判定情報には、 分類ラベルまたはスコアが含まれる。なお、GUIでは、スコアのうち特定の数値範囲を指定し、 該当するコンテンツをリスト化して表示することができる。なお、コンテンツに与えられる特徴量 は、管理パラメータ(メタ情報)である。

Description

コンテンツの分類方法および分類モデルの生成方法
 本発明の一態様は、コンピュータ装置を利用したコンテンツの分類方法、コンテンツの分類システム、分類モデルの生成方法、およびグラフィカルユーザインターフェースに関する。
 なお、本発明の一態様は、コンピュータ装置に関する。本発明の一態様は、コンピュータ装置を利用した電子化されたコンテンツ(テキストデータ、画像データ、音声データ、または動画データ)の分類方法に関する。特に、本発明の一態様は、コンテンツの集まりから機械学習を用いて効率的に分類するコンテンツの分類システムに関する。なお、本発明の一態様は、コンピュータ装置がプログラムによって管理するグラフィカルユーザインターフェースを用いたコンテンツの分類方法、コンテンツの分類システム、および分類モデルの生成方法に関する。
 ユーザは、コンテンツの集まりから、ユーザの指定するトピックスに関する情報を容易に分類し、且つ抽出したいと考えている。但し、大量のコンテンツから目的の条件に合ったコンテンツを分類する場合、個人が有する知識、または経験などによりコンテンツの分類結果にばらつきが生じる。
 近年では、個人が有する知識、経験によって分類されたコンテンツの分類結果を教師データとしてコンピュータ装置に与え、コンテンツの分類方法を機械学習させる提案がされている。例えば特許文献1では、ユーザの指定するトピックスと関連性の高いドキュメントを決定するための機械学習のアプローチが開示されている。
特開2009−104630号公報
 あるコンテンツの集合を目的に即して分類を行う場合がある。本発明の一態様では、コンテンツが、特許の場合について説明する。特許には、それぞれ固有の特許番号が与えられている。したがって以降では、コンテンツを特許番号と言い換えて説明する場合がある。なお、本発明の一態様で扱うコンテンツの分類方法では、特許番号に与えられる複数の管理パラメータに着目する。但し、コンテンツは、特許文献に限定されない。コンテンツは、テキストデータ、画像データ、音声データ、または動画データなどの情報を扱うことができる。
 コンテンツは、様々なメタ情報を用いて管理されている。メタ情報とは、あるコンテンツそのものではなく、そのコンテンツが属する属性、または関連する情報を記述するデータを示す。一例として特許番号には、コンテンツの内容として特許請求の範囲、要約書、図面、および明細書が関連付けられている。さらに特許番号には、メタ情報(評価情報、経過日数、ファミリー情報など)が与えられ、メタ情報を用いた管理がされている。特許番号は、メタ情報を用いることで重要度に応じた分類が行われる。分類の精度や効率は、対象となる文書の内容にもよるが、ユーザの経験や熟練により差が出やすく、また、大量の文書を分類する必要があり、効率化の課題があった。
 機械学習を用いた分類モデルの生成には、大量の学習データを用意する必要があるためユーザに過度の負担を強いる、という課題がある。学習データに含まれる分類されたコンテンツ数のばらつきが分類モデルの精度に影響を与える、という課題がある。
 上記問題に鑑み、本発明の一態様は、効率良く分類モデルを生成し、それを用いて情報を分類する方法を提供することを課題の一とする。または、本発明の一態様は、対話的に分類モデルを生成するためのグラフィカルユーザインターフェースを提供することを課題の一とする。または、本発明の一態様は、確率の高い情報を分類するプログラムを提供することを課題の一とする。
 なお、これらの課題の記載は、他の課題の存在を妨げるものではない。なお、本発明の一態様は、これらの課題の全てを解決する必要はないものとする。なお、これら以外の課題は、明細書、図面、請求項などの記載から、自ずと明らかとなるものであり、明細書、図面、請求項などの記載から、これら以外の課題を抽出することが可能である。
 コンピュータ装置が有する記憶装置には、プログラムが保存されている。プログラムは、コンピュータ装置が有する表示装置にグラフィカルユーザインターフェース(以下、GUI)を介して様々な情報を表示することができる。なお、ユーザは、GUIを介してコンピュータ装置にプログラムの操作、情報の付与、データベースに対する応答、機械学習の指示などを行えるものとする。また、プログラムは、GUIを介して機械学習による演算処理結果、データベースから学習用コンテンツまたは未分類のコンテンツをダウンロードした内容などを表示装置に表示することができる。なお、以降において、単にコンテンツと記す場合は、学習用コンテンツ、未分類のコンテンツ、または分類済コンテンツを含む。
 提案するコンテンツの分類システムは、機械学習を利用してコンテンツの分類モデルを生成し、生成されるコンテンツの分類モデルを用いて、未分類のコンテンツを分類する。例えば、複数のメタ情報を有するコンテンツを学習用コンテンツとする。学習用コンテンツは、さらに学習用ラベルが付与されることで学習用コンテンツから特徴ベクトルを生成する。なお特徴ベクトルを生成する場合、メタ情報または学習用ラベルは、学習用コンテンツの特徴量として扱うことができる。
 学習用コンテンツは、教師データとして扱われる。分類モデルは、学習用コンテンツを基に機械学習を行うことで取得することができる。ここで得られる分類モデルは、複数のメタ情報を有するコンテンツの分類を行う。なお、分類の種類は、ユーザの目的により2種類でも良いし、3種類以上としても良い。ユーザは、分類モデルを利用することで、全ての文書を手作業または目視で判断するよりも短い時間で文書全体の分類を行うことが可能となる。
 なお、学習用コンテンツは、データベースに保存されている学習用コンテンツからダウンロードすることができる。もしくは、コンピュータ装置の記憶装置に保存されている学習用コンテンツを用いることができる。なお、学習用コンテンツには、学習用ラベルが含まれて管理されていてもよい。さらに、データベースに保存されている分類モデルをダウンロードしてもよい。もしくは、コンピュータ装置の記憶装置に保存されている分類モデルを用いてもよい。
 本発明の一態様は、学習用コンテンツと、コンテンツと、を有し、学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、コンテンツには、第2の特徴量が付与される。複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、複数の第1の分類モデルを用いて第2の分類モデルを生成するステップと、第2の分類モデルを用いて複数のコンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、を含むコンテンツの分類方法である。
 本発明の一態様は、学習用コンテンツと、コンテンツと、を有し、学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、コンテンツには、第2の特徴量が付与される。複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、複数の第1の分類モデルの出力から平均値を算出するステップと、複数の平均値を用いて第2の分類モデルを生成するステップと、第2の分類モデルを用いて複数のコンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、を含むコンテンツの分類方法である。
 本発明の一態様は、学習用コンテンツと、コンテンツと、を有し、学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、コンテンツには、第2の特徴量が付与される。複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、複数の第1の分類モデルがそれぞれ第一の評価基準による評価をするステップと、複数の第1の分類モデルがそれぞれ第二の評価基準による評価をするステップと、複数の第一の評価基準による評価結果と、第二の評価基準と、による評価結果から第2の分類モデルを生成するステップと、第2の分類モデルを用いて複数のコンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、を含むコンテンツの分類方法である。
 上記構成において、第一の評価基準は精度(Precision)であり、第二の評価基準は感度(Sensitivity)であるコンテンツの分類方法が好ましい。
 上記各構成において、任意の学習用コンテンツを用いて第1の分類モデルを生成するステップを含むコンテンツの分類方法が好ましい。
 上記各構成において、学習用コンテンツには、さらに分類情報が与えられ、第2の分類モデルの出力を用いて分類ラベルが付与された複数のコンテンツから分類情報と同じ判定情報を有するコンテンツを選択してグラフィカルユーザインターフェース内に表示するステップと、を含むコンテンツの分類方法が好ましい。
 上記各構成において、学習用コンテンツ又コンテンツに与えられる特徴量は、管理パラメータであるコンテンツの分類方法が好ましい。
 上記各構成において、判定情報には、分類ラベルまたはスコアを含むコンテンツの分類方法が好ましい。
 上記構成において、グラフィカルユーザインターフェースは、スコアのうち特定の数値範囲を指定し、該当するコンテンツをリストとして表示するステップを含むコンテンツの分類方法が好ましい。
 本発明の一態様は、精度良く情報を分類する方法を提供することができる。または、本発明の一態様は、精度良く情報を分類するユーザインターフェースを提供することができる。または、本発明の一態様は、精度良く情報を分類するプログラムを提供することができる。
 また、本発明の一態様は、機械学習を利用した分類モデルを生成するための対話的なインターフェースをユーザに提供することができ、教師データの用意や学習結果の評価といったユーザの負担を軽減することができる。
 なお本発明の一態様の効果は、上記列挙した効果に限定されない。上記列挙した効果は、他の効果の存在を妨げるものではない。なお他の効果は、以下の記載で述べる、本項目で言及していない効果である。本項目で言及していない効果は、当業者であれば明細書または図面等の記載から導き出せるものであり、これらの記載から適宜抽出することができる。なお、本発明の一態様は、上記列挙した効果、および/または他の効果のうち、少なくとも一つの効果を有するものである。したがって本発明の一態様は、場合によっては、上記列挙した効果を有さない場合もある。
図1は、分類方法を説明するフローチャートである。
図2は、分類方法を説明するフローチャートである。
図3は、分類システム100と、ネットワークとの接続について説明する図である。
図4は、分類システムを説明するブロック図である。
図5A、図5Bは、グラフィカルユーザインターフェースを説明する図である。
図6は、分類モデルの生成方法を説明する図である。
図7は、分類モデルの生成方法を説明する図である。
図8は、分類モデルの生成方法を説明する図である。
図9は、グラフィカルユーザインターフェースを説明する図である。
図10は、グラフィカルユーザインターフェースを説明する図である。
 本実施の形態では、コンテンツの分類方法について図1乃至図10を用いて説明する。
 本実施の形態で説明するコンテンツの分類方法は、コンピュータ装置上で動作するプログラムによって制御される。プログラムは、コンピュータ装置が有するメモリ、またはストレージに保存されている。もしくは、ネットワーク(LAN(Local Area Network)、WAN(Wide Area Network)、インターネットなど)を介して接続されているコンピュータ、またはデータベースを有するサーバコンピュータに保存されている。
 なお、コンピュータ装置が有する表示装置には、ユーザがプログラムに与えるデータと、当該データをコンピュータ装置が有する演算装置によって演算した結果を表示することができる。なお、装置の構成に関しては、図4を用いて詳細に説明をする。
 表示装置に表示されるデータは、例えばリスト化された表示形式に従うことで、ユーザが認識しやすくなり、操作性が向上する。したがってユーザが、表示装置を介してコンピュータ装置が有するプログラムと簡便にやり取りするためのインターフェースをGUIとして説明する。
 ユーザは、GUIを介してプログラムが有するコンテンツの分類方法を利用することができる。ユーザは、GUIを用いてコンテンツの分類操作を簡便にすることができる。また、ユーザは、GUIを介することでコンテンツの分類結果を視覚的に判断しやすくなる。また、ユーザは、GUIを介することでプログラムを簡便に操作することができる。なお、当該コンテンツとは、テキストデータ、画像データ、音声データ、または動画データなどの情報を示している。
 次に、GUIを用いたコンテンツの分類方法を、GUIの操作手順に従い説明する。最初に、データ処理部について説明する。データ処理部は、データ収集部とデータ生成部を有する。例えば、データ収集部は、GUIを介してデータベースから複数のコンテンツからなるファイルを取得する。また、データ生成部は、ユーザがGUIを介してコンテンツに対して学習用ラベルを付与することで学習用コンテンツを生成することができる。もしくは、学習用ラベルが付与された学習用コンテンツをデータベースから取得してもよい。なお、複数のコンテンツとは、コンピュータ装置が有するメモリまたはストレージに保存されているファイル、もしくはネットワークに接続されたデータベース、コンピュータ、またはデータサーバなどに保存されているデータである。
 したがって、当該データベースには、複数の学習用コンテンツまたは複数の未分類のコンテンツがリスト化されて保存されていることが好ましい。なお、学習用コンテンツおよび未分類のコンテンツには、複数の特徴量と学習用ラベルが付与されている。当該学習用ラベルは、GUIを介してユーザが修正することができる。当該学習用ラベルを学習用コンテンツに付与した場合、学習用ラベルが付与された学習用コンテンツは、データベースに保存することができる。
 学習用コンテンツには、学習用ラベルが付与されない検証用コンテンツを含むことができる。検証用コンテンツは、学習用コンテンツを用いて生成する分類モデルを検証するために用いることができる。
 一例として、コンテンツが、特許番号の場合について説明する。特許番号には、特許番号の特徴量として複数のメタ情報が付与される。一例として、メタ情報には、評価情報、経過日数、ファミリーの数、ファミリーの状態、出願種別、ライフ、ファミリー内のペンディング数、ファミリー内の放棄数、費用、発明者数、分野、またはクレーム数などがある。つまり、メタ情報とは、コンテンツの管理パラメータである。なお、ファミリーには、特許ファミリー、またはパテントファミリーなどが含まれる。
 次に、学習処理部について説明する。学習処理部では、学習用コンテンツを用いて分類モデルを生成するステップを有する。学習処理部は、分類モデル生成部、または分類モデル評価部を有する。
 分類モデル生成部は、分類モデルを生成することができる。分類モデル生成部は、複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップを含み、当該複数の第1の分類モデルを用いて第2の分類モデルを生成するステップを含む。第1の分類モデルまたは第2の分類モデルの出力値をGUIに表示することができる。ユーザは、当該出力値に対する第1の分類モデルの学習用ラベルを付与(修正を含む)することができる。もしくは、ユーザは、当該出力値に対する新しい学習用コンテンツを追加することができる。
 分類モデル評価部は、検証用コンテンツを用いて分類モデル生成部で生成される分類モデルを評価する。当該分類モデルを用いて検証用コンテンツを推論する場合、当該分類モデルは、推論した結果を判定情報として出力する。GUIは、それぞれの評価用コンテンツに判定情報を付与して表示することができる。
 なお、ユーザは、分類モデル評価部の出力結果を判定し、必要に応じて学習用ラベルを修正し、分類モデル生成部にて分類モデルを更新することができる。または、学習用コンテンツを、追加することで、分類モデル生成部にて分類モデルを更新することができる。
 次に判定処理部について説明する。判定処理部では、分類推論部とリスト生成部を有する。例えば、分類推論部は、分類モデル生成部で生成される第1の学習モデルと第2の学習モデルを用いて複数の未分類のコンテンツを推論し分類する。当該分類モデルは、それぞれのコンテンツに対し推論結果を判定情報として付与する。
 リスト生成部は、当該判定情報が与えられたコンテンツからユーザの求める形式のリストを生成しGUIに表示することができる。例えば、それぞれのコンテンツが出願国ごとに管理されている場合、出願国を分類情報とすることができる。出願国を分類情報とする場合、生成される分類モデルは、出願国ごとに異なる分類モデルを生成することが好ましい。なお分類情報は、出願国に限定されない。例えば、コンテンツが有するメタ情報の一つを分類情報とすることができる。
 メタ情報を分類情報とする場合について説明する。例えば、特許ファミリーの状態を分類情報としてもよい。特許番号には、親出願の特許番号、分割出願の特許番号などがメタ情報で与えられている場合がある。親出願の特許番号から分割出願が可能な状態、親出願の特許番号から分割出願が不可能な状態、親出願の特許番号が権利持続中の状態、親出願の特許番号が権利喪失の状態、分割出願の特許番号からさらに分割出願が可能な状態、分割出願の特許番号からさらに分割出願が不可能な状態、分割出願の特許番号が権利持続中の状態、分割出願の特許番号が権利喪失の状態などによって、異なる分類モデルを生成し、当該分類モデルを用いてそれぞれ推論することができる。
 つまり、判定処理部は、分類モデルを用いて複数の未分類のコンテンツを推論することができる。推論される結果は、それぞれのコンテンツに判定情報として付与しGUIに表示するステップを含む。判定情報には、少なくとも分類ラベルと、スコア(確率)と、が含まれる。また、GUIは、スコアのうち特定の数値範囲を指定し、該当するコンテンツをリストとして表示するステップが含まれる。
 上述した分類モデル生成部とは異なる例を示す。分類モデル生成部は、複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、当該複数の第1の分類モデルの出力から平均値を算出するステップと、複数の当該平均値を用いて第2の分類モデルを生成するステップと、を含む。なお、第1の分類モデルまたは第2の分類モデルの出力値をGUIに表示することができる。ユーザは、当該出力値に対して、第1の分類モデルの学習用ラベルを修正することができる。もしくは、ユーザは、当該出力値に対して学習用コンテンツを追加することができる。なお、当該平均値とは、相加平均計算、相乗平均計算、または調和平均計算のいずれか一を用いて算出することを意味する。
 第2の分類モデルは、複数の当該平均値を用いて生成される。第2の分類モデルでは、第1の分類モデルの出力が平均化されるため学習用コンテンツが有する外れ値などのノイズ成分の影響を低減することができる。
 次に、上述した分類モデル生成部とは異なる例を示す。分類モデル生成部は、複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、複数の第1の分類モデルがそれぞれ第一の評価基準による評価をするステップと、複数の第1の分類モデルがそれぞれ第二の評価基準による評価をするステップと、複数の第一の評価基準による評価結果と第二の評価基準による評価結果から第2の分類モデルを生成するステップと、を含む。なお、第1の分類モデルまたは第2の分類モデルの出力値は、GUIに表示することができる。ユーザは、当該出力値に対して、第1の分類モデルの学習用ラベルを修正することができる。もしくは、ユーザは、当該出力値に対して学習用コンテンツを追加することができる。なお、第一の評価基準は混同行列の精度であり、第二の評価基準は混同行列の感度である。
 複数の第1の分類モデルの出力を、第一の評価基準および第二の評価基準により評価した結果を用いて第2の分類モデルを生成する。なお、第一の評価基準では、混同行列の精度を学習用ラベルに対する適合率と言い換えることができる。第二の評価基準では、混同行列の感度を学習用ラベルに対する再現率と言い換えることができる。よって生成される分類モデルは、複数の第1の分類モデルの当該適合率および当該再現率を含むことができる。当該複数の第1の分類モデルを用いて生成する第2の分類モデルは、分類精度が向上する。
 上述した分類モデル生成部において、例えば、生成されたm個(mは自然数を表す)の第1の分類モデルを用いて第2の分類モデルを生成してもよい。
 なお、上述した分類モデル生成部がk個(kは自然数を表す)の学習用コンテンツを有する場合、第1の分類モデルは、k個以下の任意の学習用コンテンツを生成することができる。また、学習用コンテンツがk個選択される場合、任意の学習用コンテンツは、異なる2つの学習用モデルを含むことができる。また、k個の異なる番号が付与された学習用コンテンツは、番号によってソートされる順に従いq個(qは自然数を表す)ずつを用いて第1の分類モデルを生成することができる。
 プログラムは、データベースから読み込んだコンテンツをGUIに表示することができる。コンテンツは、リスト化されたメタ情報を有することが好ましい。GUIは、GUIが有する表示形式に従い当該コンテンツを表示する。なお、コンテンツに付与されるリスト化されたメタ情報は、レコードと呼ばれる単位で管理されることが好ましい。例えば、それぞれのレコードは、番号に紐つけられたID(Identification)、コンテンツ(画像データ、音声データ、または動画データ)、またはメタ情報などによって構成される。
 本明細書では、メタ情報に着目して機械学習し、当該機械学習によって分類モデルを生成する。分類モデルは、メタ情報を解析して特徴ベクトル化されたコンテンツを分類する。
 また、上述したコンテンツの分類方法では、学習用ラベルを教師データとして用いない機械学習による分類を行うことができる。例えば、分類モデルには、K−means、またはDBSCAN(density−based spatial clustering of applications with noise)などのアルゴリズムを用いることができる。
 また、プログラムは、複数のメタ情報および学習用ラベルが付与された学習用コンテンツを用いて機械学習を用いて分類モデルを生成することができる。分類モデルには、決定木、ナイーブベイズ、KNN(k Nearest Neighbor)、SVM(Support Vector Machines)、パーセプトロン、ロジスティック回帰、またはニューラルネットワークなどのアルゴリズムを用いることができる。
 さらに、プログラムは、学習用コンテンツの数に応じて分類モデルを切り替えることができる。例えば、学習用コンテンツの数が少ないときは決定木、ナイーブベイズ、ロジスティック回帰、学習用コンテンツの数が一定量以上あればSVM、ランダムフォレスト、ニューラルネットワークを用いるとしてもよい。なお、本実施の形態で使用する分類モデルは、決定木のアルゴリズムの一つであるランダムフォレストを用いている。さらに、メタ情報の選択方法、学習用コンテンツの選択方法、または第1の分類モデルの選択方法としてランダムサンプリングまたはクロスバリエーションを用いることができる。もしくは、付与された番号のソートされる順に従いq個ずつを選択することができる。
 続いて、図面を用いてコンテンツの分類方法を説明する。図1は、本実施の一態様であるコンテンツの分類方法を説明するフローチャートである。なお、コンテンツの分類方法は、コンピュータ装置上で動作するプログラムによって制御される。したがって、プログラムが、データ処理部、学習処理部、または判定処理部を有することでコンテンツを分類することができる。プログラムは、GUIを介してユーザの求めるコンテンツの分類をすることができる。つまり上述したそれぞれの処理部で処理される内容は、プログラムのステップに相当する。
 ステップS11では、ユーザがGUIを介してコンテンツが含まれるファイルをロードする指示をすることができる。当該ファイルは、データ処理部が有するデータベースに保存されている。なお、ファイルには、学習用コンテンツ、未分類のコンテンツなどが含まれる。
 したがって、当該データベースには、複数の学習用コンテンツ、または複数の未分類のコンテンツがリスト化されて保存されていることが好ましい。ユーザは、GUIに表示されている学習用コンテンツの学習用ラベルを付与もしくは修正することができる。なお、ファイルには、学習用ラベルが付与されていない検証用コンテンツを含むことができる。
 ステップS12は、ロードしたファイルを用いて分類モデルを生成する学習処理部である。生成した分類モデルは、検証用コンテンツを評価し、評価結果をGUIに表示することができる。ユーザは、当該評価結果に対し、学習用ラベルの修正、学習用コンテンツの追加などを指示することができる。
 なお、ユーザは、学習用コンテンツの有するメタ情報の経時的変化を予測し、メタ情報を更新することができる。ユーザがメタ情報を更新する場合、分類モデルは、分類モデルの経時的な変化を含むことができる。したがって、ユーザは、コンテンツの経時的な分類の変化を得ることができる。分類モデルは、価値が増大していくと予想されるコンテンツ群、または価値が減少していくと予想されるコンテンツ群を分類することができる。
 ステップS13は、判定処理部である。ステップS12で生成される分類モデルを用いて未分類のコンテンツを推論する。当該分類モデルは、未分類のコンテンツに対し推論結果をもとに判定情報を付与することができる。判定処理部では、当該判定情報が与えられたコンテンツをユーザの求める形式でGUIに表示することができる。当該判定情報には、少なくとも分類ラベルと、スコアと、が含まれる。また、GUIは、スコアのうち特定の数値範囲を指定し、該当するコンテンツを表示することができる。
 続いて、図2では、図1のフローチャートをより詳細に説明する。まず、ステップS11の詳細について説明する。ステップS11のデータ処理部は、ステップS21のデータ収集部およびステップS22のデータ生成部を有する。
 ステップS21のデータ収集部について説明する。ステップS21のデータ収集部は、データベースからファイルをロードすることができる。なお、メタ情報またはコンテンツなどは、異なるデータベースによって管理することができる。メタ情報は、コンテンツを扱う会社、組織、またはユーザにおいて異なる場合がある。したがって、データ収集部は、異なるデータベースからコンテンツに関するメタ情報を集める機能を有する。なお、それぞれのデータベースは、異なる建物、異なる地域、または異なる国に設置することができる。
 次にステップS22のデータ生成部について説明する。データ生成部は、レコードと呼ばれる単位でコンテンツとメタ情報を管理することができる。例えば、それぞれのレコードは、番号に紐つけられたID、コンテンツ(画像データ、音声データ、または動画データ)、またはメタ情報などによって構成される。また、ユーザは、GUIに表示されるコンテンツに対して学習用ラベルを付与することで、学習用コンテンツを生成することができる。
 続いて、ステップS12の詳細について説明する。ステップS12の学習処理部は、ステップS23の分類モデル生成部、ステップS24の分類モデル評価部、およびステップS25の出力結果判定処理を有する。
 ステップS23の分類モデル生成部について説明する。分類モデル生成部は、コンテンツの分類モデルを生成することができる。分類モデル生成部は、複数の学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成することができる。当該複数の第1の分類モデルを用いて第2の分類モデルを生成することができる。GUIは、第1の分類モデルまたは第2の分類モデルの出力値を表示することができる。
 なお、後述するステップS25の出力結果判定処理後、ユーザは、当該出力値に対する第1の分類モデルの学習用ラベルを付与(修正を含む)することができる。もしくは、ユーザは、当該出力値に対する新しい学習用コンテンツを追加することができる。学習用コンテンツが有するメタ情報の経時的変化を予測し、メタ情報を更新することができる。なお、ユーザによるメタ情報を更新することで得られる効果は、ステップS12の説明を参酌することができる。
 次にステップS24の分類モデル評価部について説明する。分類モデル評価部は、検証用コンテンツを用いることで、分類モデル生成部で生成される分類モデルを評価することができる。当該分類モデルは、検証用コンテンツを推論した結果を判定情報として出力する。GUIは、それぞれの評価用コンテンツに判定情報を付与して表示することができる。
 次にステップS25の出力結果判定処理について説明する。例えば、ユーザは、ステップS24の分類モデル評価部の出力結果を判断し、コンテンツの分類モデルが十分に学習したと判定することができる。ユーザは、GUIに対し分類モデルの生成完了(OK)の指示を与える。例えば、ユーザがコンテンツの分類モデルの学習が不十分(NG)と判定することができる。ユーザは、ステップS23に戻り学習用ラベルの変更、学習用コンテンツの追加、またはメタ情報の更新などにより分類モデルを更新する。
 続いて、ステップS13の詳細について説明する。ステップS13の判定処理部は、ステップS26の分類推論部、およびステップS27のリスト作成部を有する。
 ステップS26の分類推論部について説明する。分類推論部は、分類モデル生成部で生成される第1の学習モデルと第2の学習モデルを用いて複数の未分類のコンテンツを推論し分類する。なお、分類推論部にはステップS22のデータ生成部で生成された未分類のコンテンツが与えられる。当該分類モデルは、それぞれのコンテンツに対し推論結果を判定情報として付与する。
 ステップS27のリスト作成部について説明する。リスト生成部は、当該判定情報が与えられたコンテンツをユーザの求める形式にリスト化しGUIに表示することができる。なお、それぞれのコンテンツには、メタ情報とは異なる分類情報が与えられてもよい。例えば、分類情報が学習用コンテンツに与えられている場合、生成される分類モデルは、分類情報ごとに異なる分類モデルを生成することができる。もしくは、コンテンツが有するメタ情報の一つを分類情報とすることができる。
 なお、当該判定情報には、少なくとも分類ラベルと、スコアとが含まれる。また、GUIは、スコアのうち特定の数値範囲を指定し、該当するコンテンツをリスト化してGUIに表示することができる。
 図3は、上述したコンテンツの分類方法を有する分類システム100と、ネットワーク(NetWork)との接続について説明する図である。
 分類システム100は、通信網LAN1と接続される。通信網LAN1には、データベースDB1、またはクライアントコンピュータCL1乃至CLn(nは自然数である)などが接続されている。また、通信網LAN1は、ネットワークを介して通信網LAN2と接続することができる。なお、ネットワークは、インターネット、通信網WAN、もしくは衛星通信を用いることができる。通信網LAN2には、データベースDB2、またはクライアントコンピュータCL11乃至CL1nなどが接続されている。
 分類システム100は、データベースDB1、データベースDB2、クライアントコンピュータCL1乃至CLn、またはクライアントコンピュータCL11乃至CL1nに記憶されているコンテンツを含むファイルを用いて、コンテンツの生成、コンテンツの分類、モデル生成、且つ未分類のコンテンツを分類することができる。
 また、ユーザは、分類システム100で動作するプログラムから、GUIに対して指示を出すことができる。例えば、ユーザは、インターネットを介して異なる国に設置されたデータベースの情報を用いて上述した分類モデルを生成し、未分類のコンテンツを分類することができる。つまり、コンテンツまたはメタ情報は、異なるデータベースまたはクライアントコンピュータに記憶されていてもよい。
 なお、GUIは、データベースDB1、データベースDB2、クライアントコンピュータCL1乃至CLn、またはクライアントコンピュータCL11乃至CL1nのコンピュータ装置の記憶装置に記憶された分類システム100によって分類された分類結果を表示することができる。
 図4は、図3で説明した分類システム100を説明するブロック図である。分類システム100は、GUI(Graphical User Interface)110、演算部120、および記憶部130を有している。GUI110は、入力部111および出力部112を有している。入力部111は、コンテンツのロード元を選択する機能と、学習用ラベルを入力する機能とを有する。出力部112は、データベースなどからロードしたコンテンツリストを表示する機能と、分類モデルが出力する判定情報を表示する機能とを有する。なお表示されるコンテンツに含まれるメタ情報は、GUIを介してユーザが修正することができる。
 演算部120は、データ処理部121、学習処理部122、および判定処理部123を有している。データ処理部121は、データ収集部およびデータ生成部を有している。学習処理部122は、分類モデルを作成する分類モデル生成部、および分類モデルを分類する分類モデル評価部を有している。なお、分類モデル評価部の出力結果は、ユーザによる判定が行われる評価結果判定処理の機能を有している。判定処理部123は、分類推論部と、分類推論部によって分類された結果をリスト化する出力リスト作成部を有する。演算部120は、コンピュータ装置が有する記憶部に記憶されたプログラムがマイクロプロセッサを用いて演算処理する。ただし、プログラムは、DSP(Digtal signal Processor)、またはGPU(Graphics Processing Unit)を用いて演算処理することができる。
 記憶部130には、データベースなどからロードして生成したコンテンツおよびメタ情報がリスト化されて一時的に記憶される。
 記憶部130は、例えば、1T(トランジスタ)1C(容量)型のメモリセルを備えたDRAM(ダイナミックランダムアクセスメモリ)を用いることができる。なお、DRAMのメモリセルに用いられるトランジスタは、OSトランジスタを用いてもよい。OSトランジスタは、半導体層に金属酸化物を有するトランジスタである。メモリセルにOSトランジスタが用いられるメモリ装置を「OSメモリ」と呼ぶ。ここでは、OSメモリの一例として、1T1C型のメモリセルを有するRAMのことを、「DOSRAM(Dynamic Oxide Semiconductor RAM)」と呼ぶ。
 OSトランジスタのオフ電流は非常に小さい。したがってDOSRAMは、リフレッシュの頻度を低減できるためリフレッシュ動作に要する電力を削減できる。ここでいう、オフ電流とは、トランジスタがオフ状態の場合、ソースとドレインとの間に流れる電流をいう。トランジスタがnチャネル型である場合、例えば、しきい値電圧が0V乃至2V程度であれば、ゲートとソース間の電圧が負の電圧であるときのソースとドレインとの間に流れる電流をオフ電流と呼ぶことができる。
 図5Aは、GUI30の構成を説明する図である。GUI30は、一例として、p個の学習用コンテンツをリスト表示する管理画面を示す。学習用コンテンツは、レコード単位で管理される。当該レコードには、番号(No)31、コンテンツ(ID)32、特徴量を示すメタ情報(Feature)33(メタ情報(F1)33a乃至メタ情報(Fm)33m)、分類情報(Case)34(分類情報(C1)34a乃至分類情報(Cq)34q)、学習用ラベル(J−Label)35などが含まれる。一例として、図5Aでは、学習用ラベル35が、“Yes”、“No”の2値のいずれかの値を与えるが、学習用ラベル35は、2値に限定されず、3値以上でもよい。
 図5Bは、GUI30Aの構成を説明する図である。GUI30Aは、評価推論部にてn個の未分類のコンテンツを推論し、当該推論した判定情報をリスト表示する管理画面を示す。未分類のコンテンツは、学習用コンテンツと同じく、番号31、コンテンツ32、メタ情報33、分類情報34を有する。さらに、それぞれのレコードには、分類ラベル(A−Label)36、スコア(Score)37が判定情報として付与される。
 なお、GUI30およびGUI30Aは、同じ表示画面で管理することができる。後述する図9または図10では、学習用コンテンツと判定情報とを同じ管理画面に表示することができるGUIの表示例を示す。
 図6は、機械学習によって上述した学習用コンテンツSampleに関連付けられた複数の特徴Featureを用いて分類モデルを生成する方法を説明する図である。特徴Featureは、それぞれがメタ情報のいずれか一を示し、コンテンツを管理するための管理パラメータに相当する。なお、本実施の一態様では、演算部F、演算部S、演算部V、第1の分類モデル、および第2の分類モデルを用いて分類モデルの生成方法を説明する。
 学習用コンテンツSample(1)乃至学習用コンテンツSample(k)は、それぞれj個の特徴Feature、および学習用ラベルLabelが付与されている。一例として、演算部F1は、学習用コンテンツSample(1)からコンピュータが処理できる形式の特徴ベクトルVlabel1(1)を生成することができる。また、演算部Fkは、学習用コンテンツSample(k)からコンピュータが処理できる形式の特徴ベクトルVlabel1(k)を生成することができる。なお、特徴ベクトルVlabel1(1)は、演算部F1がそれぞれの特徴に対して異なる重み係数を与えることで生成することができる。また、特徴ベクトルVlabel1(1)は、ランダムに選択されるj個以下の特徴Featureを用いて生成することができる。
 次に、複数の第1の分類モデルを生成する。演算部S1乃至演算部Smは、それぞれ異なる第1の分類モデルに相当する。一例として、演算部S1が特徴ベクトルVlabel1(1)乃至特徴ベクトルVlabel1(k)を用いて第1の分類モデルを生成することができる。なお、演算部S1に与えられる特徴ベクトルVlabel1はk個以下であればよい。異なる例として、演算部Smが異なる特徴ベクトルVlabel1(1)乃至特徴ベクトルVlabel1(k)を用いて第1の分類モデルを生成することができる。よって、異なる2つの第1の分類モデルは、それぞれk個以下の特徴ベクトルVlabel1を含み、且つ、特徴ベクトルVlabel1のいずれか一が同じ特徴ベクトルを含むことができる。
 第1の分類モデルを生成するために選択されるk個の学習用コンテンツSampleは、ランダムに選択されてもよいし、学習用コンテンツに付与された番号でソートされた順番に選択してもよい。学習用コンテンツがランダムに選択される場合、第1の分類モデルは学習用コンテンツのばらつきを含むことができる。また、学習用コンテンツに付与された番号でソートされた順番に選択される場合、時系列、またはメタ情報のいずれか一の特徴に基づいて付与された番号に沿った傾向を含むことができる。
 したがって、第1の分類モデルは、学習用コンテンツSample(1)乃至学習用コンテンツSample(k)から生成される特徴ベクトルVlabel1を用いて特徴ベクトルVlabel2を生成することができる。
 第2の分類モデルは、演算部V1によって生成される。一例として、演算部V1は、m個の特徴ベクトルVlabel2を用いて第2の分類モデルを生成するステップを有する。なお、第2の分類モデルは、特徴ベクトルVlabel2(1)乃至特徴ベクトルVlabel2(m)を用いることで異なる特徴を有する分類モデルを生成することができる。
 したがって、第2の分類モデルは、学習用コンテンツSample(1)乃至学習用コンテンツSample(k)から生成される特徴ベクトルVlabel1を用いて出力値POUTを出力することができる。GUIは、出力値POUTを表示することができる。なお、出力値POUTには、判定情報である分類ラベルと、スコアとが含まれる。したがって、第2の分類モデルは、コンテンツの分類を行うことができる。また第2の分類モデルは、それぞれのコンテンツに判定情報を付与することができる。
 図6で説明した当該分類モデルを用いて推論するには、当該分類モデルの学習用コンテンツSampleに未分類のコンテンツを与えることで判定結果を得る。ただし、未分類のコンテンツには、学習用コンテンツと異なり学習用ラベルが付与されていない。
 図7は、図6と異なる分類モデルを生成する方法を説明する図である。図7では、図6と異なる点について説明し、発明の構成(または実施例の構成)において、同一部分または同様な機能を有する部分には同一の符号を異なる図面間で共通して用い、その繰り返しの説明は省略する。
 図7では、m個の特徴ベクトルVlabel2の平均値Avを算出し、特徴ベクトルVlabel_aを生成する。第2の分類モデルは、p個の特徴ベクトルVlabel_aを用いて生成することができる。第2の分類モデルは、m個の特徴ベクトルVlabel2の平均値Avを算出することで、異なる特徴を有する分類モデルを生成することができる。生成される分類モデルは、コンテンツの分類を正確にすることができる。
 図8は、図7とは異なる分類モデルを生成する方法を説明する図である。図8では、図7と異なる点について説明し、発明の構成(または実施例の構成)において、同一部分または同様な機能を有する部分には同一の符号を異なる図面間で共通して用い、その繰り返しの説明は省略する。
 図8では、m個の特徴ベクトルVlabel2を評価するための評価基準が評価判定部JGに与えられる。例えば、評価判定部JG1には、第一の評価基準として精度(Precision)が与えられそれぞれの特徴ベクトルVlabel2(1)を評価することができる。次に、評価判定部JG1は、第二の評価基準として感度(Sensitivity)が与えられそれぞれの特徴ベクトルVlabel2(1)を評価することができる。評価判定部JG1は、評価結果Vlabel_b(1)を出力する
 第2の分類モデルは、評価結果Vlabel_b(1)乃至評価結果Vlabel_b(p)を用いて生成される。例えば、複数の特徴ベクトルVlabel2は、異なる第1の評価基準および第2の評価基準によって評価されてもよいし、同じ評価基準で評価されてもよい。なお、図8では記載していないが、図7と同様に、第1の評価基準、および第2の評価基準による評価結果Vlabel_bの平均値を算出することができる。
 第2の分類モデルは、m個の特徴ベクトルVlabel2の評価結果を用いることで、異なる特徴を有する分類モデルを生成することができる。生成される分類モデルは、コンテンツの分類を正確にすることができる。
 図9は、GUI50を説明する図である。GUI50は、コンテンツ(学習用コンテンツ、未分類のコンテンツ、分類済のコンテンツ)の表示領域、コンテンツを含むファイルのダウンロード元を選択するアイコン58a、選択されたファイルが保存されるアドレス情報を表示するテキストボックス58b、機械学習を実行するアイコン(Learning Start)59を有している。
 表示領域には、一例として8個のレコードがロードされている例を示している。それぞれのレコードは、番号(No)51、ID(Index)52、特徴量(Feature)53、分類情報(Case)54、学習用ラベル(JL)55、分類ラベル(AL)56、およびスコア(Prob)57の構成要素を有する。特徴量53は、詳細な情報として特徴量F(1)53a乃至特徴量F(j)53jを表示することができる。なお、jは自然数である。また、分類情報54は、詳細な情報として分類情報C(1)54a乃至分類情報C(4)54dを表示することができる。なお、分類情報は、自然数で表すことができる種類を有することができる。
 なお、図9は、分類モデルによって学習用コンテンツおよび未分類のコンテンツが分類された結果をGUI50に表示する例を示している。
 一例として、レコード番号のNo1乃至No3は、学習用コンテンツに相当する。学習用コンテンツには、学習用ラベルが与えられ、且つ、レコード番号No1乃至No3には、分類情報が与えられている。
 レコード番号のNo4乃至No8は、分類済のコンテンツに相当する。分類済のコンテンツには、分類ラベル56、スコア57が与えられている。なお、図9では、レコード番号No1およびレコード番号No3を学習することで得られる分類モデルを用いて、レコード番号No4乃至レコード番号No7を分類した結果を表示している。一例として、レコード番号No2を学習することで得られた分類モデルを用いて、レコード番号No8を分類した結果を表示している。図9ではスペースの関係上レコード数は8までしか表示していないが、レコード数は複数の種類を扱うことができる。
 但し、レコード数を大量に扱う場合、表示上の課題がある。したがって、分類ラベル56またはスコア57はソート機能を備えることが好ましい。ソート条件の一例として、GUIは、分類ラベル56が“Yes”の判定結果を選択して表示することができる。また、GUIは、スコア57の数値範囲を指定して表示することができる。GUIに上述したソート条件を与える場合、GUIは、教師データを与えた学習コンテンツと同じような特徴を有するコンテンツを分類し表示することができる。
 例えば、コンテンツが特許番号の場合について説明する。特許番号には、複数のメタ情報が付与されている。特許の権利が維持されている特許番号の場合、学習用ラベルには、“Yes”を与える。特許の権利が放棄されている特許番号の場合、学習用ラベルには、“No”を与える。続いて機械学習を実行し分類モデルを生成する。
 上述した分類モデルは、未分類のコンテンツに対し、判定情報を付与することができる。判定情報には、分類ラベル56と、スコア57が表示される。一例として、ユーザは、ソート機能を用いて分類ラベル56に“No”を与える。さらに、スコア57には、“0.8”乃至“1.0”を設定する。GUIに上述したソート条件を与えることで、GUIは、特許番号が放棄された学習用コンテンツと同じような特徴を有するレコードを選択して表示することができる。
 図10は、図9とは異なるGUI50Aを説明する図である。図10は、レコード数を大量に扱う場合の効率的なGUIの表示例である。なお、図10では、図9と異なる点について説明し、発明の構成(または実施例の構成)において、同一部分または同様な機能を有する部分には同一の符号を異なる図面間で共通して用い、その繰り返しの説明は省略する。
 図10では、任意の選択した分類情報に関するレコードごと分類して表示することができる点が図9と異なっている。図10では、分類情報C(1)乃至分類情報C(4)の種類に応じて表示を切替ることができる。
 ユーザが、レコードに付与された複数の特徴量53、および分類モデルの判定情報を見て、十分な分類精度が得られていれば分類モデルの更新を終了する。ユーザが、レコードに付与された判定情報を見て、分類精度が十分でないと判断した場合は、ユーザ指定ラベルが付与されていないレコードに対して学習用ラベルを付与し、アイコン59を押すことで、分類モデルの更新を行うことができる。なお、学習用コンテンツが有するメタ情報の経時的変化を予測し、特徴量53を更新してもよい。ユーザが特徴量53を更新する場合、分類モデルは、分類モデルの経時的な変化を含むことができる。したがって、ユーザは、コンテンツの経時的な分類の変化を得ることができる。分類モデルは、価値が増大していくと予想されるコンテンツ群、または価値が減少していくと予想されるコンテンツ群を分類していくことができるようになる。
 図示していないが、特徴量53、分類情報54a乃至分類情報54d、学習用ラベル55、分類ラベル56、またはスコア57に含まれる数値やラベル情報は、表示させる並び順を変更することができる、もしくは、選択した数値やラベル情報は、フィルタ機能を用いて必要な並び順にソートし表示することができる。これにより、ユーザは、分類モデルの判定結果を効率良く評価することができる。
 図1乃至図10を用いて説明したコンテンツの分類方法は、確率の高い情報を分類する方法を提供することができる。例えば、GUIは、確率の高い情報を分類するのに適している。プログラムは、分類モデルに新しい教師データ(学習用ラベル)が与えられることで分類モデルを更新することができる。プログラムは、分類モデルが更新されることで確率の高い情報を分類することができる。
 さらに、生成した分類モデルは電子機器本体または外部のメモリに保存することができ、新たなファイルの分類の際に呼び出して使うことができる。さらに、新たな教師データを追加しながら上記で説明した方法に従って分類モデルを更新することができる。
 以上、本実施の形態で示す構成、方法は、他の実施の形態で示す構成、方法と適宜組み合わせて用いることができる。
CL1:クライアントコンピュータ、CL1n:クライアントコンピュータ、CL11:クライアントコンピュータ、CLn:クライアントコンピュータ、DB1:データベース、DB2:データベース、LAN1:通信網、LAN2:通信網、Vlabel1:特徴ベクトル、Vlabel2:特徴ベクトル、31:番号、32:コンテンツ、33:メタ情報、34:分類情報、35:学習用ラベル、50:GUI、50A:GUI、51:番号、53:特徴量、54:分類情報、56:分類ラベル、57:スコア、58a:アイコン、58b:テキストボックス、59:アイコン、100:分類システム、110:GUI、111:入力部、112:出力部、120:演算部、121:データ処理部、122:学習処理部、123:判定処理部、130:記憶部

Claims (9)

  1.  学習用コンテンツと、コンテンツと、を有し、
     前記学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、
     前記コンテンツには、第2の特徴量が付与され、
     複数の前記学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、
     前記複数の第1の分類モデルを用いて第2の分類モデルを生成するステップと、
     前記第2の分類モデルを用いて複数の前記コンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、
     を含むコンテンツの分類方法。
  2.  学習用コンテンツと、コンテンツと、を有し、
     前記学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、
     前記コンテンツには、第2の特徴量が付与され、
     複数の前記学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、
     前記複数の第1の分類モデルの出力から平均値を算出するステップと、
     複数の前記平均値を用いて第2の分類モデルを生成するステップと、
     前記第2の分類モデルを用いて複数の前記コンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、
     を含むコンテンツの分類方法。
  3.  学習用コンテンツと、コンテンツと、を有し、
     前記学習用コンテンツには、第1の特徴量および学習用ラベルが付与され、
     前記コンテンツには、第2の特徴量が付与され、
     複数の前記学習用コンテンツを用いて複数の第1の分類モデルを機械学習によって生成するステップと、
     前記複数の第1の分類モデルがそれぞれ第一の評価基準による評価をするステップと、
     前記複数の第1の分類モデルがそれぞれ第二の評価基準による評価をするステップと、
     複数の前記第一の評価基準による評価結果と、前記第二の評価基準と、による評価結果から第2の分類モデルを生成するステップと、
     前記第2の分類モデルを用いて複数の前記コンテンツに判定情報を付与しグラフィカルユーザインターフェース内に表示するステップと、
     を含むコンテンツの分類方法。
  4.  請求項3において、
     前記第一の評価基準は精度であり、
     前記第二の評価基準は感度であるコンテンツの分類方法。
  5.  請求項1乃至3のいずれかにおいて、
     任意の前記学習用コンテンツを用いて前記第1の分類モデルを生成するステップを含むコンテンツの分類方法。
  6.  請求項1乃至3のいずれかにおいて、
     前記学習用コンテンツには、さらに分類情報が与えられ、
     前記第2の分類モデルの出力を用いて分類ラベルが付与された複数の前記コンテンツから前記分類情報と同じ前記判定情報を有するコンテンツを選択して前記グラフィカルユーザインターフェース内に表示するステップと、
     を含むコンテンツの分類方法。
  7.  請求項1乃至3のいずれかにおいて、
     前記学習用コンテンツ又前記コンテンツに与えられる特徴量は、管理パラメータであるコンテンツの分類方法。
  8.  請求項1乃至3のいずれかにおいて、
     前記判定情報には、分類ラベルまたはスコアを含むコンテンツの分類方法。
  9.  請求項8において、
     前記グラフィカルユーザインターフェースは、前記スコアのうち特定の数値範囲を指定し、該当するコンテンツをリストとして表示するステップを含むコンテンツの分類方法。
PCT/IB2019/060377 2018-12-13 2019-12-03 コンテンツの分類方法および分類モデルの生成方法 Ceased WO2020121115A1 (ja)

Priority Applications (7)

Application Number Priority Date Filing Date Title
US17/311,730 US20220027799A1 (en) 2018-12-13 2019-12-03 Content classification method and classification model generation method
DE112019006203.4T DE112019006203T5 (de) 2018-12-13 2019-12-03 Verfahren zur Klassifizierung von Inhalten und Verfahren zur Erzeugung eines Klassifizierungsmodells
KR1020217016573A KR20210100613A (ko) 2018-12-13 2019-12-03 콘텐츠의 분류 방법 및 분류 모델의 생성 방법
CN201980078452.2A CN113168421A (zh) 2018-12-13 2019-12-03 内容的分类方法及分类模型的生成方法
JP2020559054A JP7730641B2 (ja) 2018-12-13 2019-12-03 コンテンツの分類方法
JP2024227415A JP7734819B2 (ja) 2018-12-13 2024-12-24 コンテンツの分類システム
JP2025140625A JP2025161942A (ja) 2018-12-13 2025-08-26 コンテンツの分類方法

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2018-233037 2018-12-13
JP2018233037 2018-12-13

Publications (1)

Publication Number Publication Date
WO2020121115A1 true WO2020121115A1 (ja) 2020-06-18

Family

ID=71075466

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/IB2019/060377 Ceased WO2020121115A1 (ja) 2018-12-13 2019-12-03 コンテンツの分類方法および分類モデルの生成方法

Country Status (6)

Country Link
US (1) US20220027799A1 (ja)
JP (3) JP7730641B2 (ja)
KR (1) KR20210100613A (ja)
CN (1) CN113168421A (ja)
DE (1) DE112019006203T5 (ja)
WO (1) WO2020121115A1 (ja)

Cited By (1)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2023539240A (ja) * 2020-08-25 2023-09-13 アルテリックス インコーポレイテッド ハイブリッド機械学習

Families Citing this family (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
CN113841137A (zh) 2019-05-24 2021-12-24 株式会社半导体能源研究所 文件检索系统及文件检索方法
US11501304B2 (en) * 2020-03-11 2022-11-15 Synchrony Bank Systems and methods for classifying imbalanced data
WO2024006154A1 (en) * 2022-06-28 2024-01-04 HonorEd Technologies, Inc. Comprehension modeling and ai-sourced student content recommendations

Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008242880A (ja) * 2007-03-28 2008-10-09 Kenwood Corp コンテンツ表示システム、コンテンツ表示方法および車載用の情報端末装置
JP2013161298A (ja) * 2012-02-06 2013-08-19 Nippon Steel & Sumitomo Metal 分類器作成装置、分類器作成方法、及びコンピュータプログラム
WO2014203328A1 (ja) * 2013-06-18 2014-12-24 株式会社日立製作所 音声データ検索システム、音声データ検索方法、及びコンピュータ読み取り可能な記憶媒体

Family Cites Families (9)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7287012B2 (en) 2004-01-09 2007-10-23 Microsoft Corporation Machine-learned approach to determining document relevance for search over large electronic collections of documents
US20150206069A1 (en) * 2014-01-17 2015-07-23 Matthew BEERS Machine learning-based patent quality metric
US10331782B2 (en) * 2014-11-19 2019-06-25 Lexisnexis, A Division Of Reed Elsevier Inc. Systems and methods for automatic identification of potential material facts in documents
US20170316438A1 (en) * 2016-04-29 2017-11-02 Genesys Telecommunications Laboratories, Inc. Customer experience analytics
JP6280997B1 (ja) * 2016-10-31 2018-02-14 株式会社Preferred Networks 疾患の罹患判定装置、疾患の罹患判定方法、疾患の特徴抽出装置及び疾患の特徴抽出方法
US20180197087A1 (en) * 2017-01-06 2018-07-12 Accenture Global Solutions Limited Systems and methods for retraining a classification model
US11323564B2 (en) * 2018-01-04 2022-05-03 Dell Products L.P. Case management virtual assistant to enable predictive outputs
US20190279073A1 (en) * 2018-03-07 2019-09-12 Sap Se Computer Generated Determination of Patentability
US10740621B2 (en) * 2018-06-30 2020-08-11 Microsoft Technology Licensing, Llc Standalone video classification

Patent Citations (3)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008242880A (ja) * 2007-03-28 2008-10-09 Kenwood Corp コンテンツ表示システム、コンテンツ表示方法および車載用の情報端末装置
JP2013161298A (ja) * 2012-02-06 2013-08-19 Nippon Steel & Sumitomo Metal 分類器作成装置、分類器作成方法、及びコンピュータプログラム
WO2014203328A1 (ja) * 2013-06-18 2014-12-24 株式会社日立製作所 音声データ検索システム、音声データ検索方法、及びコンピュータ読み取り可能な記憶媒体

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2023539240A (ja) * 2020-08-25 2023-09-13 アルテリックス インコーポレイテッド ハイブリッド機械学習
JP7806027B2 (ja) 2020-08-25 2026-01-26 アルテリックス インコーポレイテッド ハイブリッド機械学習

Also Published As

Publication number Publication date
JP2025161942A (ja) 2025-10-24
DE112019006203T5 (de) 2021-09-02
JPWO2020121115A1 (ja) 2020-06-18
JP2025041828A (ja) 2025-03-26
KR20210100613A (ko) 2021-08-17
US20220027799A1 (en) 2022-01-27
JP7730641B2 (ja) 2025-08-28
CN113168421A (zh) 2021-07-23
JP7734819B2 (ja) 2025-09-05

Similar Documents

Publication Publication Date Title
JP7734819B2 (ja) コンテンツの分類システム
US11900320B2 (en) Utilizing machine learning models for identifying a subject of a query, a context for the subject, and a workflow
US20220253725A1 (en) Machine learning model for entity resolution
CN115062732B (zh) 基于大数据用户标签信息的资源共享合作推荐方法及系统
CN111930518B (zh) 面向知识图谱表示学习的分布式框架构建方法
CN107562789A (zh) 知识库问题更新方法、客服机器人以及可读存储介质
Xie et al. Factorization machine based service recommendation on heterogeneous information networks
CN115082102B (zh) 能源消费者的预测性细分
CN110705719A (zh) 执行自动机器学习的方法和装置
WO2019015631A1 (zh) 生成机器学习样本的组合特征的方法及系统
Nagaraju et al. Boost customer churn prediction in the insurance industry using meta-heuristic models
US11323564B2 (en) Case management virtual assistant to enable predictive outputs
Shaji et al. Weather prediction using machine learning algorithms
Olorunnimbe et al. Dynamic adaptation of online ensembles for drifting data streams
AU2021258019A1 (en) Utilizing machine learning models to generate initiative plans
Wang et al. Explaining genetic programming-evolved routing policies for uncertain capacitated arc routing problems
Ming-Te et al. Using data mining technique to perform the performance assessment of lean service
Obulaporam et al. GCRITICPA: A CRITIC and grey relational analysis based service ranking approach for cloud service selection
de Sá et al. Algorithm recommendation for data streams
JP2022069602A (ja) 営業活動支援システム、営業活動支援方法および営業活動支援プログラム
Derbel et al. Automatic classification and analysis of multiple-criteria decision making
CN117038074B (zh) 基于大数据的用户管理方法、装置、设备及存储介质
JP6942672B2 (ja) 情報処理装置、情報処理方法、及び情報処理プログラム
Chen et al. Using granular computing model to induce scheduling knowledge in dynamic manufacturing environments
Das et al. Decision Rule Prediction for Assessment of Rotor Spun Cotton Yarn Strength Using Rough Set

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 19895325

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2020559054

Country of ref document: JP

Kind code of ref document: A

ENP Entry into the national phase

Ref document number: 20217016573

Country of ref document: KR

Kind code of ref document: A

122 Ep: pct application non-entry in european phase

Ref document number: 19895325

Country of ref document: EP

Kind code of ref document: A1