WO2015045155A1 - コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム - Google Patents
コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム Download PDFInfo
- Publication number
- WO2015045155A1 WO2015045155A1 PCT/JP2013/076545 JP2013076545W WO2015045155A1 WO 2015045155 A1 WO2015045155 A1 WO 2015045155A1 JP 2013076545 W JP2013076545 W JP 2013076545W WO 2015045155 A1 WO2015045155 A1 WO 2015045155A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- attribute
- reference word
- corpus
- unit
- data
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Images
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
- G06F40/103—Formatting, i.e. changing of presentation of documents
- G06F40/117—Tagging; Marking up; Designating a block; Setting of attributes
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/10—Text processing
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/237—Lexical tools
- G06F40/242—Dictionaries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/253—Grammatical analysis; Style critique
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/30—Semantic analysis
Definitions
- One embodiment of the present invention relates to a corpus generation device, a corpus generation method, and a corpus generation program.
- Patent Document 1 For example, a technique described in Patent Document 1 is known as a technique for automatically attaching a tag to a natural sentence such as a description of a product.
- Patent Document 1 refers to dictionary data, extracts general-purpose specific expressions and semantic role words from an input document, estimates the sentence structure of the input document based on general-purpose specific expressions and semantic role words, There is described a sentence processing apparatus that assigns a semantic tag to an input document in accordance with rules defined in advance for the sentence structure.
- a corpus generation device includes a web page acquisition unit that acquires a web page including explanatory text data related to a presentation target, and a reference word acquisition unit that acquires a reference word that is an attribute value related to the presentation target from the web page And a broader term that is higher than the reference word acquired by the reference word acquisition unit is extracted from the storage unit that stores the hierarchical relationship information representing the vertical relationship between the attribute values, and the explanatory word data includes the broader word.
- an assigning unit that assigns an attribute tag corresponding to the reference word to a broader term included in the explanatory text data
- an output unit that outputs the explanatory text data provided with the attribute tag by the assigning unit as corpus data And comprising.
- a corpus generation method includes a web page acquisition step of acquiring a web page including explanatory text data related to a presentation target, and a reference word acquisition step of acquiring a reference word that is an attribute value related to the presentation target from the web page And a broader term that is higher than the reference word acquired in the reference word acquisition step is extracted from the storage unit that stores the hierarchical relationship information representing the vertical relationship between the attribute values, and the explanatory word data includes the broader word.
- an adding step for assigning an attribute tag corresponding to the reference word to a broader term included in the explanatory text data
- an output step for outputting the explanatory text data to which the attribute tag is given in the assigning step as corpus data And including.
- a corpus generation program provides a computer, a web page acquisition unit that acquires a web page including explanatory text data related to a presentation target, and a reference that acquires a reference word that is an attribute value related to the presentation target from the web page
- a broader word that is higher than the reference word acquired by the reference word acquisition unit is extracted from the word acquisition unit and the storage unit that stores the hierarchical relationship information indicating the vertical relationship between the attribute values, and the broader word is described in the explanatory text data. Is included, the attribute part corresponding to the reference word is added to the broader terms included in the explanatory text data, and the explanatory text data to which the attribute tag is assigned is output as corpus data Function as an output unit.
- the attribute tag corresponding to the reference word is assigned to the broader word included in the explanatory text data.
- the reference word is an attribute value related to the presentation target
- the broader term is also an attribute value related to the presentation target.
- by assigning an attribute tag to a broader term in the explanatory text data it is possible to suppress erroneous assignment of the attribute tag to a phrase that does not indicate the feature to be presented. Then, by outputting the explanatory text data to be presented to which the attribute tag is assigned in this way as corpus data, high-quality corpus data can be generated.
- the adding unit when the reference word is included in the explanatory text data, the adding unit further adds an attribute tag corresponding to the reference word to the reference word included in the explanatory text data. Also good. According to such a form, an accurate attribute tag can be assigned to more attribute values to be presented in the explanatory text data.
- the web page may further include an attribute list in which attribute names and attribute values related to the presentation target are associated, and the reference word acquisition unit may acquire the attribute value in the attribute list as a reference word.
- the attribute value in the attribute list is highly likely to be an attribute value related to the presentation target. Therefore, according to this embodiment, an appropriate reference word can be acquired.
- the reference word acquisition unit may use a syntax analyzer to search the explanatory text data for a highly probable phrase that is an attribute value related to the presentation target, and acquire the searched word as a reference word. According to such a form, a reference word can be acquired even from a web page that does not include an attribute list.
- the hierarchical relation information has a tree structure with the attribute value as a node, and the assigning unit searches the partial tree with the reference word as a root node and does not have a branch from the hierarchical relation information.
- the reference word is extracted with respect to the one or more attribute values included in the explanatory text data.
- An attribute tag corresponding to may be further added.
- the hierarchical relationship information hierarchically represents the vertical relationship between attribute values
- the assigning unit is a case where a plurality of attribute values exist in a hierarchy immediately below the reference word in the hierarchical relationship information
- an attribute tag corresponding to a reference word is further added to the one attribute value included in the explanatory text data. May be.
- the attribute value is related to the presentation target. There is a high probability of being an attribute value. Therefore, according to this embodiment, an accurate attribute tag can be assigned to more attribute values to be presented in the explanatory text data.
- the information processing apparatus may further include a presentation target information registration unit that stores a set of attribute names and attribute values related to the presentation target acquired from the corpus data output from the output unit in association with the web page. According to such a form, a combination of an attribute name and an attribute value can be used as a clue for searching for a presentation target.
- a corpus generation device includes a corpus acquisition unit that acquires corpus data in which an attribute tag is assigned to an attribute value, and a reference that acquires a reference word that is an attribute value related to a presentation target from the corpus data.
- a broader term that is higher than the reference word acquired by the reference word acquisition unit is extracted from the word acquisition unit and a storage unit that stores the hierarchical relationship information representing the vertical relationship between the attribute values, and the attribute is different from the broader word
- the removal unit that removes the attribute tag attached to the attribute value different from the broader term from the corpus data, and the output that outputs the corpus data from which the attribute tag is removed by the removal unit A section.
- a corpus generation method includes a corpus acquisition step of acquiring corpus data in which an attribute tag is assigned to an attribute value, and a reference for acquiring a reference word that is an attribute value related to a presentation target from the corpus data
- a broader term that is higher than the reference word acquired in the reference word acquisition step is extracted from the storage unit that stores the hierarchical relationship information representing the hierarchical relationship between the word acquisition step and the attribute value, and an attribute different from the broader word
- a removal step for removing the attribute tag attached to the attribute value different from the broader term from the corpus data, and an output for outputting the corpus data from which the attribute tag is removed in the removal step Steps.
- a corpus generation program provides a computer with a corpus acquisition unit that acquires corpus data in which an attribute tag is assigned to an attribute value, and a reference word that is an attribute value related to a presentation target from the corpus data.
- a broader term that is higher than the reference word acquired by the reference word acquisition unit is extracted from the reference word acquisition unit to be acquired and a storage unit that stores the hierarchical relationship information representing the vertical relationship between the attribute values.
- a removal unit that removes the attribute tag given to the attribute value different from the broader term from the corpus data, and corpus data from which the attribute tag is removed by the removal unit It functions as an output unit that outputs.
- the attribute tag given to the attribute value different from the broader word is removed from the corpus data.
- the reference word is an attribute value related to the presentation target
- the broader term is also an attribute value related to the presentation target.
- the attribute value different from the broader term is not the attribute value related to the presentation target. For this reason, corpus data to which an attribute tag is accurately assigned can be generated by removing an attribute tag assigned to an attribute value different from the broader term included in the corpus data.
- a corpus generation device includes a web page acquisition unit that acquires a web page including explanatory text data related to a presentation target, and dictionary data storage that stores dictionary data in which attribute names and attribute values are associated with each other.
- a candidate extraction unit that extracts candidates for attribute values related to the presentation target included in the description data
- a reference word acquisition unit that acquires a reference word that is an attribute value related to the presentation target from the web page
- the attribute value From the storage unit that stores the hierarchical relationship information representing the hierarchical relationship of the above, a broader term that is higher than the reference word acquired by the reference word acquisition unit is extracted, and among the attribute value candidates included in the explanatory text data,
- An attribute unit that assigns an attribute tag corresponding to an attribute value candidate to an attribute value candidate that matches the word, and outputs explanatory text data to which the attribute tag is assigned by the assigning unit as corpus data And an output section that, a.
- a corpus generation method includes a web page acquisition step of acquiring a web page including explanatory text data related to a presentation target, and dictionary data storage that stores dictionary data in which attribute names and attribute values are associated with each other.
- a candidate extraction step for extracting attribute value candidates related to the presentation target included in the description data a reference word acquisition step for acquiring a reference word that is an attribute value related to the presentation target from the web page, and the attribute value From the storage unit that stores the hierarchical relationship information representing the hierarchical relationship of the above, a broader term that is higher than the reference word acquired in the reference word acquisition step is extracted, and among the attribute value candidates included in the explanatory text data,
- An assigning unit that assigns an attribute tag corresponding to an attribute value candidate to an attribute value candidate that matches the word, and an explanation that the attribute tag is assigned in the assigning step Comprising an output step of outputting the data as a corpus data.
- a corpus generation program stores a computer, a web page acquisition unit that acquires a web page including explanatory text data related to a presentation target, and dictionary data in which attribute names and attribute values are associated with each other.
- a candidate extraction unit that refers to the dictionary data storage unit and extracts candidate attribute values related to the presentation target included in the explanatory text data;
- a reference word acquisition unit that acquires a reference word that is an attribute value related to the presentation target from the web page; From the storage unit that stores the hierarchical relationship information representing the hierarchical relationship between the attribute values, a broader term that is higher than the reference word acquired by the reference word acquisition unit is extracted, and candidate attribute values included in the explanatory text data are extracted.
- the assigning unit that assigns the attribute tag corresponding to the attribute value candidate to the attribute value candidate that matches the broader term, and the explanatory text data to which the attribute tag is assigned by the assigning unit are copied.
- An output unit for outputting as the path data, and to function.
- the attribute tag corresponding to the reference word is given to the attribute value that matches the broader word among the candidate attribute values included in the explanatory text data.
- the reference word is an attribute value related to the presentation target
- the broader term is also an attribute value related to the presentation target.
- an attribute value different from the reference word and its broader term is not an attribute value related to the presentation target.
- the attribute tag corresponding to the attribute value candidate is assigned to the attribute value candidate that matches the broader word. It is possible to prevent an attribute tag from being erroneously assigned to a word that is not a value.
- FIG. It is a figure which shows an example of up-and-down relationship information.
- (A) is an example of the explanatory text data contained in a goods page
- (b) is an example of the explanatory text data to which the attribute tag was provided.
- (A) is an example of the explanatory text data contained in a goods page
- (b) is an example of the explanatory text data to which the attribute tag was provided.
- FIG. 1 is a diagram illustrating a functional configuration of a corpus generation device 1 according to the first embodiment.
- the corpus generation device 1 includes a web page acquisition unit 12, a reference word acquisition unit 14, a grant unit 16, an output unit 18, and a product information registration unit (presentation target information registration unit) 20.
- the corpus generation device 1 is a device that assigns an attribute tag to explanatory text data included in a web page related to a presentation target, and outputs the explanatory text provided with the attribute tag as corpus data.
- the presentation object in the web page used for generation of corpus data is goods is explained, the presentation object is not limited to goods.
- the corpus generation device 1 includes a CPU 101 that executes an operating system, application programs, and the like, a main storage unit 102 that includes a ROM and a RAM, and an auxiliary storage unit 103 that includes a hard disk.
- Each function of the corpus generation device 1 shown in FIG. 1 reads predetermined software on the CPU 101 and the main storage unit 102, and controls the communication control unit 104, the input unit 105, the output unit 106, and the like under the control of the CPU 101.
- the operation is realized by reading and writing data in the main storage unit 102 and the auxiliary storage unit 103. Data and databases necessary for processing are stored in the main storage unit 102 and the auxiliary storage unit 103.
- the web page acquisition unit 12 is a functional element that acquires a product page (web page) including explanatory text data related to a product.
- the web page acquisition unit 12 can access the product page storage unit 22 and acquires the product page stored in the product page storage unit 22.
- the web page acquisition unit 12 outputs the acquired product page to the reference word acquisition unit 14.
- the product page storage unit 22 is a functional element that stores product pages related to products sold on the electronic commerce site.
- FIG. 3 is an example of a product page M1 stored in the product page storage unit 22.
- the product page M1 includes a product image B1, explanation text data D1, and an attribute list L1.
- the description sentence data D1 is character string data composed of natural sentences explaining the characteristics of the products sold on the product page M1.
- the attribute list L1 is an attribute name A1 associated with a product and an attribute value V1 associated with each other.
- the attribute name is an item name of the attribute
- the attribute value is a specific attribute value.
- the attribute list L1 is created by, for example, an operator of a store that sells products.
- the reference word acquisition unit 14 is a functional element that acquires a reference word from the web page acquisition unit 12.
- the reference word is a highly probable phrase that is an attribute value that represents the feature of the product, and is used as a reference for extracting the attribute value included in the description sentence data D1.
- the reference word acquisition unit 14 may acquire the attribute value V1 as a reference word from the attribute list L1 included in the product page output from the web page acquisition unit 12.
- the attribute value V1 included in the attribute list L1 is information registered by an operator of a store that sells products. Therefore, the attribute value V1 is a phrase that is very likely to be an attribute value that represents the characteristics of the product sold on the product page M1. is there.
- the reference word acquisition unit 14 may acquire a reference word not from the attribute list L1 included in the product page M1 but from the explanatory text data D1 in the product page M1.
- the reference word acquisition unit 14 extracts, as a reference word, an attribute value that has a high probability of being an attribute value representing the characteristics of the product from among the attribute values included in the explanatory text data D1 using a syntax analyzer.
- the reference word acquisition unit 14 searches the description sentence data D1 for sentences and phrases in which attribute names and attribute values appear successively, and acquires the attribute values as reference words.
- a sentence “the place of production is Burgundy” is detected in the explanatory text data D1
- “burgundy” is detected as a reference word.
- the reference word acquisition unit 14 represents an attribute, for example, “product” of “French” or “Chateau” of “Chateau Margaux”, and explained using a prefix and a suffix added to the attribute value.
- a reference word may be acquired from sentence data.
- the reference word acquisition unit 14 outputs information indicating the acquired reference word to the assigning unit 16.
- the assigning unit 16 extracts a broader term that is an attribute value belonging to a higher hierarchy than the reference word acquired by the reference word acquiring unit 14 from the hierarchical relationship information stored in the hierarchical relationship information storage unit 24, and provides explanatory text data It is a functional element that gives an attribute tag corresponding to the reference word to the broader term included in D1.
- the assigning unit 16 can access the vertical relationship information storage unit 24 and the tag information storage unit 26.
- the hierarchical relationship information storage unit 24 stores hierarchical relationship information that hierarchically represents the conceptual vertical relationship between attribute values.
- the hierarchical relationship here is a concept including an upper-lower relationship and a partial overall relationship.
- the upper limit relation information has a tree structure having a plurality of hierarchies and having attribute values as nodes.
- the hierarchical relation information stores attribute values that conceptually represent small categories as the hierarchy deepens.
- FIG. 4 is an example conceptually showing the data structure of the hierarchical relationship information. In the upper limit relation data shown in FIG. 4, the attribute value “France” is stored in the highest hierarchy. In the lower hierarchy, attribute values “Burgundy”, “Bordeaux”, “Champagne”, etc., which are subdivisions of “France” are stored.
- attribute value “Course de Nuit”, which is a subdivision of the attribute value “burgundy”, is stored in the lower hierarchy of the attribute value “burgundy”, and the attribute value “Marsane” is stored in the lower hierarchy.
- attribute values “Saint Testev” and “Poyac”, which are sub-classes of “Bordeaux”, are stored in a lower hierarchy of the attribute value “Bordeaux”.
- the attribute value “Italy” is stored in the highest hierarchy as a series independent of the series related to the attribute value “France”.
- attribute values “Piedmont”, “Tuscany”, and the like indicating the subdivision of the attribute value “Italy” are stored.
- the tag information storage unit 26 stores tag information in which attribute names and attribute values are associated.
- FIG. 5 is a diagram illustrating an example of tag information. As shown in FIG. 5, the tag information does not have a hierarchical structure unlike the hierarchical relationship information, and a plurality of attribute values are associated with one attribute name.
- “France, Italy, Burgundy, Côte de Nuit” and the like are associated with the attribute name “production center” as attribute values.
- This tag information indicates that, for example, the phrases “France”, “Italy”, “Burgundy”, and “Côte de Nuits” are attribute values representing “origin”.
- corresponding attribute values are also associated with the attribute names “type” and “product type”.
- FIG. 6A is a diagram illustrating an example of the explanatory text data D2 included in the product page.
- the explanatory note data D2 includes “France”, “Burgundy”, and “Piedmont” as phrases relating to the production area.
- FIG. 6B is a diagram illustrating an example of the attribute list L2 included in the product page.
- the attribute values included in the attribute list L2 are acquired by the reference word acquisition unit 14 described above and output to the assigning unit 16. In the following description, it is assumed that the reference word acquisition unit 14 outputs “burgundy” that is the attribute value V2 included in the attribute list L2 to the adding unit 16 as a reference word.
- the assigning unit 16 searches for the reference word “burgundy” from the hierarchical relation information, and extracts broader terms belonging to a higher hierarchy than “burgundy”.
- the attribute value “Burgundy” is included in the layer one level below the highest layer, and therefore the assigning unit 16 places “France” belonging to the upper layer of “Burgundy” as the upper level. Extract as a word.
- the assigning unit 16 acquires an attribute name corresponding to the reference word “burgundy” with reference to the tag information stored in the tag information storage unit 26.
- the attribute name “production center” is associated with the reference word “burgundy”.
- the assigning unit 16 assigns an attribute tag indicating the attribute name corresponding to the reference word to the broader word included in the explanatory note data D2.
- the assigning unit 16 may further add an attribute tag corresponding to the reference word to the reference word included in the explanatory text data D2.
- the assigning unit 16 refers to the tag information storage unit 26 and assigns the attribute name “production center” corresponding to the reference word “burgundy” to the explanatory text data D2.
- FIG. 6C is a diagram illustrating an example of the explanatory text data C2 to which the attribute tag is assigned by the assigning unit 16.
- the explanatory text data C2 includes an attribute tag “ ⁇ Locality>... ⁇ / Locality>” indicating the attribute name “Locality” and the broader words “France” in the explanatory text data C2. It is attached to the reference word “Burgundy”.
- the assigning unit 16 does not assign an attribute tag related to the production area to the phrase “Piemonte” included in the explanatory note data D2. This is because “Piedmont” is not a broad word or a reference word, and is therefore a phrase related to the place of production, but it is assumed that it is not an attribute value indicating the characteristics of the product.
- the attribute tag indicating the attribute name is not necessarily acquired from the tag information storage unit 26.
- the assigning unit 16 may acquire the attribute name A2 corresponding to the attribute value V2 serving as the reference word in the attribute list L2 as an attribute name, and assign an attribute tag indicating the attribute name to the reference word and the broader term.
- the assigning unit 16 acquires the attribute name “production center” corresponding to the attribute value “Burgundy” in the attribute list L2. Then, the attribute tag corresponding to the attribute name “production center” is assigned to the broader term and the reference term in the explanatory text data D2.
- the assigning unit 16 outputs the explanatory text data C ⁇ b> 2 to which the attribute tag is assigned as described above to the output unit 18.
- the output unit 18 outputs the explanatory text data C2 with the attribute tag output from the assigning unit 16 to the corpus data storage unit 28 as corpus data.
- the corpus data storage unit 28 is a functional element that stores the corpus data output from the output unit 18.
- the product information registration unit 20 is a functional element that stores a product tag that is a set of attribute names and attribute values acquired from the corpus data output from the output unit 18 in association with the product page.
- the product information registration unit 20 can access the corpus data storage unit 28 and the product information storage unit 30.
- the product information storage unit 30 is a functional element that stores product information related to products sold by the virtual store.
- FIG. 7 is a diagram illustrating an example of product information.
- Each record of the product information shown in FIG. 7 includes a store ID of a virtual store that provides a product, a product ID that uniquely identifies the product, a product name, a category, a price, a URL of a product page, an inventory quantity, and a product tag. Yes.
- the information included in the product information is not limited to these.
- the merchandise information registration unit 20 acquires corpus data from the corpus data storage unit 28, and acquires a combination of an attribute name and an attribute value using an attribute tag attached to the corpus data. Then, a record of product information corresponding to the product page from which the corpus data is acquired is extracted, and a pair of an attribute name and an attribute value is registered in the product tag of the record. This product tag is used as a key for narrowing down the product page intended by the user from the electronic commerce site.
- the corpus data stored in the corpus data storage unit 28 can be constructed, for example, by using machine learning for an analysis device that automatically generates an attribute list.
- the attribute list can be automatically generated from the explanatory text data included in the product page.
- the analysis apparatus can recognize a place where an attribute value of “production center” appears in the context.
- the analysis device After applying machine learning using such corpus data CX, when the analysis device analyzes the explanatory text data D3 as shown in FIG. 8B, the analysis device has the structure and context of the sentence. From the above, it is detected that “Chile” is an attribute value representing “Origin”, and as shown in FIG. 8C, the tag “ ⁇ Origin> ... ⁇ / Origin>” is attached to the “Chile” part. Output data C3 can be generated.
- the analysis apparatus can generate an attribute list in which the attribute name “production center” and the attribute value “Chile” are associated with each other based on the output data C3. As a result, even if the product page that includes the explanatory text data D3 does not have an attribute list, the product page may include an attribute list that associates the attribute name “origin” with the attribute value “Chile”. It becomes possible.
- attribute tag adding process by the adding unit 16 can be modified in various ways. Below, the 1st modification of the provision part 16 is demonstrated.
- the assigning unit 16 searches the subtree having the reference word as the root node and having no branch from the hierarchical relation information, and obtains one or more attribute values other than the reference word included in the subtree.
- An attribute tag corresponding to the reference word is further added to one or more attribute values included in the explanatory text data.
- FIG. 9 is a diagram illustrating another example of the vertical relationship information.
- the assigning unit 16 sets “burgundy” as a root node and “burgundy”, “court de spray”, and “marsane” as nodes as subtrees having no branches.
- the containing partial tree PT is extracted.
- “Côte de Nuits” and “Marsane” are extracted by excluding the reference word “Burgundy”.
- the assigning unit 16 acquires the attribute name “production center” corresponding to the reference word “burgundy” with reference to the tag information stored in the tag information storage unit 26.
- the assigning unit 16 assigns the attribute tag “production center” to “Côtes de Nuits” and “Marsane”.
- the assigning unit 16 assigns the attribute tag “production center” only to the reference word “Bordeaux” and its broader term “France” included in the explanatory text data.
- the assigning unit 16 according to the second modification is a case where there are a plurality of attribute values in the hierarchy immediately below the reference word in the hierarchical relation information, and only one attribute value of the plurality of attribute values is described.
- an attribute tag corresponding to the reference word is further added to the one attribute value included in the explanatory sentence data.
- FIG. 10A is an example of explanatory text data D4 included in the product page.
- the descriptive text data D4 includes “France”, “Bordeaux”, and “Saint-Testef” as phrases relating to the production area.
- the assigning unit 16 searches for the attribute value belonging to the hierarchy immediately below the reference word “Bordeaux”.
- two attribute values “Saint Testev” and “Poyac” exist in the lower hierarchy of “Bordeaux”.
- the explanatory note data D4 includes only “San Testev”, which is one of the two attribute values.
- FIG. 10B shows the explanatory text data C4 in which the attribute tag is added to the explanatory text data D4 by the adding unit 16 of the present modification.
- the attribute tag “ ⁇ Locality>... ⁇ / Locality>” indicating the attribute name “Locality” has the broader words “France” in the explanatory text data C4.
- FIG. 11A is another example of the explanatory text data included in the product page.
- This descriptive text data D5 includes “France”, “Bordeaux”, “Saint Testev”, and “Poyac” as phrases relating to the production area.
- This explanatory text data D5 includes both a plurality of attribute values “Saint Testev” and “Poyac” belonging to the hierarchy immediately below the reference word “Bordeaux”. In this case, since it is difficult to determine which of the “San Testev” and “Poyac” is the attribute value related to the product, the assigning unit 16 sets “San Testev” and “Poyac”.
- FIG. 11B shows the explanatory text data C5 in which the attribute tag is added to the explanatory text data D5 by the adding unit 16 of the present modification.
- the attribute tag “ ⁇ Locality>... ⁇ / Locality>” indicating the attribute name “Locality” is the broader term “France” and the standard in the explanatory text data C4. It is attached only to the word “Bordeaux”.
- a sentence including “Saint Testev” and “poyac”, which are subordinate terms of the reference word “Bordeaux”, may not be included in the final corpus. This is to improve the quality of the corpus.
- the web page acquisition unit 12 of the corpus generation device 1 acquires a web page including descriptive text data related to a product from the product page storage unit 22 (step S1, web page acquisition step).
- the reference word acquisition part 14 acquires a reference word from the web page acquired in step S1 (step S2, reference word acquisition step).
- the reference word can be acquired from an attribute list included in the web page.
- the assigning unit 16 extracts, from the hierarchical relationship information stored in the hierarchical relationship information storage unit 24, broader terms that are attribute values belonging to a higher hierarchy than the reference word acquired in step S2 (step S3). .
- the assigning unit 16 assigns an attribute tag corresponding to the reference word to the broader term included in the description text data of the product page (step S4, granting step). At this time, an attribute tag corresponding to the reference word may be attached to the reference word included in the description text data of the product page. Further, the assigning unit 16 searches the subtree having the reference word as the root node and having no branch from the hierarchical relation information, extracts one or more attribute values other than the reference word included in the subtree, and describes the explanatory text An attribute tag corresponding to the reference word may be further added to the one or more attribute values included in the data.
- the assigning unit 16 is a case where a plurality of attribute values exist in the hierarchy immediately below the reference word in the hierarchical relationship information, and only one attribute value among the plurality of attribute values is included in the explanatory text data.
- an attribute tag corresponding to the reference word may be further added to one attribute value included in the explanatory text data.
- the output unit 18 stores the explanatory text data to which the attribute tag is assigned in step S4 in the corpus data storage unit 28 as corpus data (step S5, output step).
- the corpus generation device 1 generates corpus data.
- FIG. 13 is a diagram illustrating a configuration of the corpus generation program P1 according to the embodiment.
- the computer can be operated as the corpus generation device 1 having the above functions.
- the corpus generation program P1 includes a main module P2, a web page acquisition module P3, a reference word acquisition module P4, a grant module P5, an output module P6, and a product information registration module P7.
- the main module P2 is a part that comprehensively controls processing.
- the functions realized by executing the web page acquisition module P3, the reference word acquisition module P4, the grant module P5, the output module P6, and the product information registration module P7 are the web page acquisition unit 12 and the reference word acquisition unit, respectively. 14, the function of the granting unit 16, the output unit 18, and the product information registration unit 20.
- the corpus generation program P1 is provided after being recorded on a tangible recording medium such as a CD-ROM, DVD-ROM, or semiconductor memory.
- the corpus generation program P1 may be provided via the network N as a data signal superimposed on a carrier wave.
- the reference word when the broader word of the reference word that is the attribute value related to the presentation target is included in the explanatory text data, the reference word is compared with the broader word included in the explanatory text data. An attribute tag corresponding to is attached.
- the reference word is an attribute value related to the presentation target, it is inferred that the broader term is also an attribute value related to the presentation target. For this reason, by assigning an attribute tag to a broader term in the explanatory text data, it is possible to suppress erroneous assignment of the attribute tag to a phrase that does not indicate the feature to be presented.
- the adding unit 16 further adds an attribute tag corresponding to the reference word to the reference word included in the description sentence data, it is accurate to attribute values of more products in the description sentence data. An attribute tag can be attached.
- the reference word acquisition unit 14 acquires the attribute value in the attribute list as a reference word. Since the attribute value of the product in the attribute list is highly likely to be an attribute value related to the product, the corpus generation device 1 can acquire an appropriate reference word.
- the corpus generation device 1A is a device that removes an attribute tag assigned to an attribute value different from the broader term from corpus data to which an attribute tag has already been assigned.
- items that are different from the first embodiment will be mainly described, and description of items that are the same as or equivalent to those of the first embodiment will be omitted.
- FIG. 14 is a diagram illustrating a functional configuration of the corpus generation device 1A.
- the corpus generation device 1A includes a corpus acquisition unit 32, a reference word acquisition unit 14, a removal unit 36, an output unit 18, and a product information registration unit 20.
- the corpus acquisition unit 32 is a functional element that acquires corpus data in which an attribute tag is assigned to an attribute value from the corpus data storage unit 28.
- the corpus data acquired by the corpus acquisition unit 32 may be text data in which an attribute tag is added to an attribute value included in the explanatory text data, and includes explanatory text data and an attribute list. Web page data in which an attribute tag is assigned to an included attribute value may be used.
- the reference word acquisition unit 14 is a functional element that acquires a reference word that is an attribute value related to a product from the corpus data acquired by the corpus acquisition unit 32.
- the reference word acquisition unit 14 uses an attribute value as a reference word from the attribute list included in the corpus data. Can get.
- the corpus data acquired by the corpus acquisition unit 32 is text data in which an attribute tag is added to an attribute value included in the explanatory text data
- the reference word acquisition unit 14 uses a syntax analyzer.
- an attribute value having a high probability that is an attribute value representing the characteristics of the product can be extracted as a reference word.
- the removal unit 36 extracts a broader term that belongs to a higher rank than the reference word acquired by the reference word acquisition unit 14 from the hierarchical relationship information storage unit 24 that stores the hierarchical relationship information representing the hierarchical relationship between attribute values.
- an attribute value different from a word is included in the corpus data, it is a functional element that removes an attribute tag attached to an attribute value different from the broader word from the corpus data.
- the removing unit 36 may leave the attribute tag attached to the reference word without removing it from the corpus data.
- the removal unit 36 removes the attribute tag from the corpus data.
- the corpus data includes an attribute value different from the broader word and the reference word
- the removal unit 36 removes the attribute tag given to the attribute value different from the broader word and the reference word from the corpus data.
- FIG. 15A is a diagram illustrating an example of the corpus data C6a acquired by the corpus acquisition unit 32.
- an attribute tag “ ⁇ production area>... ⁇ / Production area>” indicating an attribute name “production area” is attached to the phrases “France”, “Burgundy”, and “Piedmont” in the corpus data C6a.
- the reference word acquisition unit 14 outputs “Burgundy” from the corpus data C6a as the reference word to the assigning unit 16.
- the removal unit 36 searches the reference word “burgundy” from the hierarchical relation information, and extracts broader terms belonging to a higher hierarchy than “burgundy”.
- the assigning unit 16 places “France” belonging to the upper hierarchy of “burgundy” as a higher rank. Extract as a word.
- the removal unit 36 acquires a word / phrase having an attribute tag attached to the corpus data C6a, and searches for a word / phrase different from the broader word.
- the removal unit 36 removes the attribute tag given to the phrase “Piedmont” included in the corpus data C6a.
- FIG. 15B is a diagram illustrating an example of corpus data C6b from which the attribute tag has been removed by the removal unit 36.
- the attribute tags given to the phrases “France” and “Burgundy” are maintained, but the attribute tags given to the phrase “Piedmont” are removed. Yes. This is because “Piedmont” is different from the broader term or the reference term, and it is assumed that it is not an attribute value indicating the feature of the product.
- the removal unit 36 outputs the corpus data from which part of the attribute tag is removed to the output unit 18.
- the output unit 18 outputs the corpus data output from the removal unit 36 to the corpus data storage unit 28 and updates the corpus data in the corpus data storage unit 28. Since the function of the merchandise information registration unit 20 is as described in the first embodiment, the description thereof is omitted.
- the removal unit 36 searches a subtree having a reference word as a root node and having no branch from the hierarchical relation information, and obtains one or more attribute values other than the reference word included in the subtree.
- the attribute tag extracted and assigned to the one or more attribute values may be left without being removed from the corpus data.
- the removal unit 36 is a case where there are a plurality of attribute values in the hierarchy immediately below the reference word in the hierarchical relation information, and only one attribute value is included in the corpus data among the plurality of attribute values.
- the attribute tag assigned to the one attribute value may be left without being removed from the corpus data.
- the corpus acquisition unit 32 of the corpus generation device 1A acquires the corpus data in which the attribute tag is given to the attribute value from the corpus data storage unit 28 (step S11, corpus acquisition step).
- the reference word acquisition unit 14 acquires a reference word from the corpus data acquired in step S1 (step S12).
- the removal unit 36 extracts a broader term that is an attribute value belonging to a higher hierarchy than the reference word acquired in step S12 from the hierarchical relationship information stored in the hierarchical relationship information storage unit 24 (step S13). .
- the removal unit 36 removes the attribute tag given to the attribute value different from the broader word from the corpus data (step S14, removal step). ). At this time, the removal unit 36 may leave the attribute tag attached to the reference word without removing it from the corpus data when an attribute value different from the reference word is included in the corpus data. Further, the removal unit 36 searches the subtree having the reference word as the root node and having no branch from the hierarchical relation information, extracts one or more attribute values other than the reference word included in the subtree, You may leave the attribute tag provided to the above attribute value, without removing it from corpus data.
- the removal unit 36 is a case where a plurality of attribute values exist in the hierarchy immediately below the reference word in the hierarchical relation information, and only one attribute value is included in the corpus data among the plurality of attribute values.
- the attribute tag assigned to the one attribute value may be left without being removed from the corpus data.
- the output unit 18 stores the corpus data from which the attribute tag has been removed in step S14 in the corpus data storage unit 28 (step S15). As described above, the corpus generation device 1A generates updated corpus data.
- the corpus generation program includes a main module, a corpus acquisition module, a reference word acquisition module, a removal module, an output module, and a product information registration module.
- the main module is a part that controls the corpus generation program centrally.
- the functions realized by executing the corpus acquisition module, reference word acquisition module, removal module, output module, and product information registration module are the corpus acquisition unit 32, reference word acquisition unit 14, removal unit 36, and output described above, respectively.
- the functions of the unit 18 and the product information registration unit 20 are the same.
- the corpus generation program according to the present embodiment can be distributed by various methods in the same manner as the corpus generation program P1 of the first embodiment.
- the attribute tag given to the attribute value different from the broader word is detected from the corpus data. Removed.
- the reference word is an attribute value related to a product
- the broader term is also an attribute value related to the product.
- an attribute value different from the reference word and its broader term is not an attribute value related to the product. For this reason, corpus data to which an attribute tag is accurately assigned can be generated by removing an attribute tag assigned to an attribute value different from the broader term included in the corpus data.
- the corpus generation device 1B extracts attribute value candidates from the web page including the description data using dictionary data, and sets the attribute value candidates that match the broader word as attribute value candidates included in the description data. This is a device for attaching a tag.
- items that are different from the first embodiment will be mainly described, and description of items that are the same as or equivalent to those of the first embodiment will be omitted.
- FIG. 17 is a diagram illustrating a functional configuration of the corpus generation device 1B. As shown in FIG. 16, in addition to the web page acquisition part 12, the reference word acquisition part 14, the provision part 16, the output part 18, and the product information registration part 20, the candidate extraction part 42 is provided.
- the candidate extraction unit 42 refers to a tag information storage unit (dictionary data storage unit) 26 that stores dictionary data in which attribute names and attribute values are associated with each other, and selects one or more attribute value candidates included in the explanatory text data. It is a functional element to be extracted.
- the candidate extraction unit 42 can access the tag information storage unit 26. As described in the first embodiment, the tag information storage unit 26 stores tag information in which attribute names and attribute values are associated. In this embodiment, this tag information also functions as dictionary data.
- the candidate extraction unit 42 extracts words / phrases that match the attribute values of the dictionary data from the description data, and sets the extracted words / phrases as candidates for attribute values related to the product. In addition, the candidate extraction unit 42 acquires attribute names corresponding to the extracted attribute value candidates from the dictionary data.
- the candidate extraction unit 42 outputs the attribute value candidates and their attribute names to the assigning unit 16.
- the assigning unit 16 extracts, from the hierarchical relation information storage unit 24, broader terms that belong higher than the reference words acquired by the reference word acquisition unit 14, and candidates for attribute values included in the explanatory sentence data Among these, an attribute tag corresponding to a candidate attribute value is assigned to a candidate attribute value that matches the broader term. On the other hand, no attribute tag is assigned to attribute value candidates that do not match the broader word among the attribute value candidates included in the explanatory text data.
- the assigning unit 16 may further add an attribute tag corresponding to the candidate attribute value to the candidate attribute value that matches the reference word among the candidate attribute values included in the explanatory text data.
- the assigning unit 16 attaches an attribute tag to the explanatory text data.
- the assigning unit 16 assigns an attribute tag corresponding to an attribute value candidate to an attribute value candidate that matches a broader word or a reference word among attribute value candidates included in the explanatory text data.
- FIG. 18A is a diagram illustrating an example of the explanatory text data D7 included in the product page.
- the explanatory note data D7 includes “France”, “Burgundy”, and “Piedmont” as terms relating to the production area.
- the reference word acquisition unit 14 acquires “Burgundy” from the product page as a reference word, and outputs this reference word to the assigning unit 16.
- the candidate extraction unit 42 refers to the dictionary data stored in the tag information storage unit 26, and extracts one or more attribute value candidates included in the explanatory text data.
- the candidate extraction unit 42 extracts the phrases “France”, “Burgundy”, and “Piedmont” included in the explanatory text data D7 as attribute value candidates.
- the candidate extraction unit 42 refers to the dictionary data and acquires the attribute name “production center” corresponding to these words / phrases.
- the assigning unit 16 searches for the reference word “burgundy” from the hierarchical relation information, and extracts broader terms belonging to a higher hierarchy than “burgundy”. In the hierarchical relation information shown in FIG. 4, since the reference word “burgundy” is included in the hierarchy one level below the highest hierarchy, the assigning unit 16 places “France” belonging to the upper hierarchy of “burgundy” as a higher rank. Extract as a word. Next, the assigning unit 16 selects an attribute value candidate that matches the broader word from the attribute value candidates.
- the assigning unit 16 assigns an attribute tag indicating the attribute name “production center” to “France” and “Burgundy” that match the broader term or the reference word.
- FIG. 18B is a diagram illustrating an example of the explanatory text data C7 to which the attribute tag is assigned by the grant unit 16.
- attribute tags are assigned to the phrases “France” and “Burgundy”, and no attribute tag is assigned to the phrase “Piedmont”. This is because “Piedmont” does not match the broader term or the reference term, so it is assumed that it is not an attribute value indicating the feature of the product.
- the assigning unit 16 searches for a subtree having a reference word as a root node and having no branch from the hierarchical relation information, and extracts one or more attribute values other than the reference word included in the subtree.
- an attribute tag corresponding to the attribute value candidate may be further added to the attribute value that matches the one or more attribute values.
- the assigning unit 16 is a case where a plurality of attribute values exist in a hierarchy immediately below the reference word in the hierarchical relationship information, and only one attribute value among the plurality of attribute values is included in the corpus data.
- the attribute value corresponding to the attribute value candidate may be further added to the attribute value matching the one attribute value among the attribute value candidates included in the explanatory note data.
- the granting unit 16 outputs the explanatory text data to which the attribute tag is given to the output unit 18.
- the output unit 18 outputs the explanatory text data output from the assigning unit 16 to the corpus data storage unit 28 as corpus data.
- the web page acquisition unit 12 of the corpus generation device 1B acquires a web page including descriptive text data about a product from the product page storage unit 22 (step S21).
- the candidate extraction unit 42 refers to the tag information storage unit 26 that stores dictionary data, and extracts one or more attribute value candidates included in the explanatory text data (step S22, candidate extraction step).
- the reference word acquisition part 14 acquires a reference word from the web page acquired in step S21 (step S23).
- the assigning unit 16 extracts, from the hierarchical relationship information stored in the hierarchical relationship information storage unit 24, broader terms that are attribute values belonging to a higher hierarchy than the reference word acquired in step S23 (step S24). .
- the assigning unit 16 assigns an attribute tag corresponding to the candidate attribute value to the attribute value that matches the broader word among the candidate attribute values included in the explanatory text data (step S25). At this time, when the attribute value that matches the reference word is included in the explanatory text data, the assigning unit 16 sets the attribute value that matches the reference word among the attribute value candidates included in the explanatory text data. An attribute tag corresponding to a value candidate may be further added.
- the assigning unit 16 searches the subtree having the reference word as a root node and having no branch from the hierarchical relation information, extracts one or more attribute values other than the reference word included in the subtree, and explains Of the attribute value candidates included in the sentence data, an attribute tag corresponding to the attribute value candidate may be further added to the attribute value candidate that matches the one or more attribute values.
- the assigning unit 16 is a case where a plurality of attribute values exist in the hierarchy immediately below the reference word in the hierarchical relationship information, and only one attribute value among the plurality of attribute values is included in the explanatory text data. In this case, an attribute tag corresponding to the attribute value candidate may be further added to the attribute value candidate that matches the attribute value among the attribute value candidates included in the explanatory text data.
- the output unit 18 stores the explanatory text data to which the attribute tag is assigned in step S25 in the corpus data storage unit 28 (step S26).
- the corpus generation device 1B generates corpus data.
- the corpus generation program includes a main module, a web page acquisition module, a reference word acquisition module, an assignment module, an output module, a product information registration module, and a candidate extraction module.
- the main module is a part that controls the corpus generation program centrally.
- the functions realized by executing the web page acquisition module, reference word acquisition module, grant module, output module, product information registration module, and candidate extraction module are the web page acquisition unit 12 and reference word acquisition unit 14 described above, respectively.
- the functions of the granting unit 16, the output unit 18, the product information registration unit 20, and the candidate extraction unit 42 are the same.
- the corpus generation program according to the present embodiment can be distributed by various methods in the same manner as the corpus generation program P1 of the first embodiment.
- the attribute tag corresponding to the reference word is assigned to the attribute value matching the broader word.
- the reference word is an attribute value related to a product
- the broader term is also an attribute value related to the product.
- an attribute value different from the reference word and its broader term is not an attribute value related to the product.
- the attribute value corresponding to the attribute value candidate is assigned to the attribute value candidate that matches the broader word among the attribute value candidates included in the explanatory text data. It is possible to prevent an attribute tag from being erroneously assigned to a word that is not.
- an attribute tag is assigned to product description data included in a wine product page.
- an attribute tag is assigned to product product description data related to a product different from wine. May be.
- the corpus generation device 1 may add attribute tags other than the production area, such as “product type”, “type”, and “capacity”, to the explanatory text data.
- the same attribute tag is assigned to the broader term, the reference word, etc. in the explanatory text data, but a different attribute tag may be assigned to each.
- a different attribute tag may be assigned to each. For example, if the broader term is "France” and the reference word is "Burgundy”, an attribute tag related to the attribute name "country” is assigned to the broader word, and an attribute tag related to the attribute name "region” is assigned to the reference word May be given.
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Health & Medical Sciences (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Artificial Intelligence (AREA)
- Document Processing Apparatus (AREA)
- Machine Translation (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
図1は、第1実施形態に係るコーパス生成装置1の機能的構成を示す図である。図1に示すように、コーパス生成装置1はウェブページ取得部12、基準語取得部14、付与部16、出力部18及び商品情報登録部(提示対象情報登録部)20を備えている。コーパス生成装置1は、提示対象に関するウェブページに含まれる説明文データに対して属性タグを付与し、属性タグが付与された説明文をコーパスデータとして出力する装置である。以下では、コーパスデータの生成に用いるウェブページにおける提示対象が商品である場合について説明するが、提示対象は商品に限定されない。
第1の変形例に係る付与部16は、上下関係情報から基準語をルートノードとし、且つ分岐を有しない部分木を検索して、部分木に含まれる基準語以外の1以上の属性値を抽出し、説明文データに含まれる1以上の属性値に対して、基準語に対応する属性タグを更に付与する。以下、図9を参照して、第1の変形例に係る付与部16の処理を具体的に説明する。図9は、上下関係情報の別の一例を示す図である。
第2の変形例に係る付与部16は、上下関係情報において基準語の直下の階層に複数の属性値が存在する場合であって、当該複数の属性値の内、一の属性値のみが説明文データに含まれる場合には、説明文データに含まれる当該一の属性値に対して、基準語に対応する属性タグを更に付与する。以下、図9~11を参照して、第2の変形例に係る付与部16に係る付与部16の処理を具体的に説明する。
次に、第2実施形態に係るコーパス生成装置1Aについて説明する。コーパス生成装置1Aは、既に属性タグが付与されたコーパスデータから上位語とは異なる属性値に対して付与されている属性タグを除去する装置である。以下では、主に第1実施形態と相違する事項について説明し、第1実施形態と同一又は同等の事項については説明を省略する。
次に、第3実施形態に係るコーパス生成装置1Bについて説明する。コーパス生成装置1Bは、説明データを含むウェブページから辞書データを用いて属性値の候補を抽出し、説明文データに含まれる属性値の候補の内、上位語と一致する属性値の候補に属性タグを付与する装置である。以下では、主に第1実施形態と相違する事項について説明し、第1実施形態と同一又は同等の事項については説明を省略する。
Claims (15)
- 提示対象に関する説明文データを含むウェブページを取得するウェブページ取得部と、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに前記上位語が含まれる場合には、前記説明文データに含まれる前記上位語に対して、前記基準語に対応する属性タグを付与する付与部と、
前記付与部により属性タグが付与された説明文データをコーパスデータとして出力する出力部と、
を備えるコーパス生成装置。 - 前記付与部は、前記説明文データに前記基準語が含まれる場合には、前記説明文データに含まれる前記基準語に対して、該基準語に対応する属性タグを更に付与する、請求項1に記載のコーパス生成装置。
- 前記ウェブページは、前記提示対象に関する属性名と属性値とが対応付けられた属性リストを更に含み、
前記基準語取得部は、前記属性リストにおける属性値を前記基準語として取得する、請求項1又は2に記載のコーパス生成装置。 - 前記基準語取得部は、構文解析器を用いて、前記説明文データから前記提示対象に関する属性値である蓋然性が高い語句を検索し、検索された語句を前記基準語として取得する、請求項1又は2に記載のコーパス生成装置。
- 前記上下関係情報は、属性値をノードとする木構造をなし、
前記付与部は、前記上下関係情報から前記基準語をルートノードとし、且つ分岐を有しない部分木を検索して、該部分木に含まれる前記基準語以外の1以上の属性値を抽出し、前記説明文データに前記1以上の属性値が含まれる場合には、前記説明文データに含まれる該1以上の属性値に対して、前記基準語に対応する属性タグを更に付与する、請求項1~4の何れか一項に記載のコーパス生成装置。 - 前記上下関係情報は、属性値間の上下関係を階層的に表しており、
前記付与部は、前記上下関係情報において前記基準語の直下の階層に複数の属性値が存在する場合であって、該複数の属性値の内、一の属性値のみが前記説明文データに含まれる場合には、前記説明文データに含まれる該一の属性値に対して、前記基準語に対応する属性タグを更に付与する、請求項1~4の何れか一項に記載のコーパス生成装置。 - 前記出力部から出力されたコーパスデータから取得される前記提示対象に関する属性名と属性値との組を前記ウェブページに関連付けて記憶する提示対象情報登録部を更に備える、請求項1~6の何れか一項に記載のコーパス生成装置。
- 提示対象に関する説明文データを含むウェブページを取得するウェブページ取得ステップと、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得ステップと、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得ステップにおいて取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに前記上位語が含まれる場合には、前記説明文データに含まれる前記上位語に対して、前記基準語に対応する属性タグを付与する付与ステップと、
前記付与ステップにおいて属性タグが付与された説明文データをコーパスデータとして出力する出力ステップと、
を含むコーパス生成方法。 - コンピュータを、
提示対象に関する説明文データを含むウェブページを取得するウェブページ取得部と、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに前記上位語が含まれる場合には、前記説明文データに含まれる前記上位語に対して、前記基準語に対応する属性タグを付与する付与部と、
前記付与部により属性タグが付与された説明文データをコーパスデータとして出力する出力部と、
して機能させる、コーパス生成プログラム。 - 属性値に対して属性タグが付与されたコーパスデータを取得するコーパス取得部と、
前記コーパスデータから提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記上位語とは異なる属性値が前記コーパスデータに含まれている場合に、前記上位語とは異なる属性値に付与された属性タグを前記コーパスデータから除去する除去部と、
前記除去部により属性タグが除去された前記コーパスデータを出力する出力部と、
を備えるコーパス生成装置。 - 属性値に対して属性タグが付与されたコーパスデータを取得するコーパス取得ステップと、
前記コーパスデータから提示対象に関する属性値である基準語を取得する基準語取得ステップと、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得ステップにおいて取得された前記基準語よりも上位に属する上位語を抽出し、前記上位語とは異なる属性値が前記コーパスデータに含まれている場合に、前記上位語とは異なる属性値に付与された属性タグを前記コーパスデータから除去する除去ステップと、
前記除去ステップにおいて属性タグが除去された前記コーパスデータを出力する出力ステップと、
を含むコーパス生成方法。 - コンピュータを、
属性値に対して属性タグが付与されたコーパスデータを取得するコーパス取得部と、
前記コーパスデータから提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記上位語とは異なる属性値が前記コーパスデータに含まれている場合に、前記上位語とは異なる属性値に付与された属性タグを前記コーパスデータから除去する除去部と、
前記除去部により属性タグが除去された前記コーパスデータを出力する出力部と、
して機能させる、コーパス生成プログラム。 - 提示対象に関する説明文データを含むウェブページを取得するウェブページ取得部と、
属性名と属性値とが関連付けられた辞書データを記憶する辞書データ記憶部を参照し、前記説明文データに含まれる前記提示対象に関する属性値の候補を抽出する候補抽出部と、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに含まれる前記属性値の候補の内、前記上位語と一致する属性値の候補に前記属性値の候補に対応する属性タグを付与する付与部と、
前記付与部により属性タグが付与された説明文データをコーパスデータとして出力する出力部と、
を備えるコーパス生成装置。 - 提示対象に関する説明文データを含むウェブページを取得するウェブページ取得ステップと、
属性名と属性値とが関連付けられた辞書データを記憶する辞書データ記憶部を参照し、前記説明文データに含まれる前記提示対象に関する属性値の候補を抽出する候補抽出ステップと、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得ステップと、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得ステップにおいて取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに含まれる前記属性値の候補の内、前記上位語と一致する属性値の候補に前記属性値の候補に対応する属性タグを付与する付与ステップと、
前記付与ステップにおいて属性タグが付与された説明文データをコーパスデータとして出力する出力ステップと、
を含むコーパス生成方法。 - コンピュータを、
提示対象に関する説明文データを含むウェブページを取得するウェブページ取得部と、
属性名と属性値とが関連付けられた辞書データを記憶する辞書データ記憶部を参照し、前記説明文データに含まれる前記提示対象に関する属性値の候補を抽出する候補抽出部と、
前記ウェブページから前記提示対象に関する属性値である基準語を取得する基準語取得部と、
属性値間の上下関係を表す上下関係情報を記憶する記憶部から、前記基準語取得部により取得された前記基準語よりも上位に属する上位語を抽出し、前記説明文データに含まれる前記属性値の候補の内、前記上位語と一致する属性値の候補に前記属性値の候補に対応する属性タグを付与する付与部と、
前記付与部により属性タグが付与された説明文データをコーパスデータとして出力する出力部と、
して機能させる、コーパス生成プログラム。
Priority Applications (3)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2014508630A JP5576003B1 (ja) | 2013-09-30 | 2013-09-30 | コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム |
| US14/420,424 US9645979B2 (en) | 2013-09-30 | 2013-09-30 | Device, method and program for generating accurate corpus data for presentation target for searching |
| PCT/JP2013/076545 WO2015045155A1 (ja) | 2013-09-30 | 2013-09-30 | コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム |
Applications Claiming Priority (1)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| PCT/JP2013/076545 WO2015045155A1 (ja) | 2013-09-30 | 2013-09-30 | コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2015045155A1 true WO2015045155A1 (ja) | 2015-04-02 |
Family
ID=51579036
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2013/076545 Ceased WO2015045155A1 (ja) | 2013-09-30 | 2013-09-30 | コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム |
Country Status (3)
| Country | Link |
|---|---|
| US (1) | US9645979B2 (ja) |
| JP (1) | JP5576003B1 (ja) |
| WO (1) | WO2015045155A1 (ja) |
Cited By (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2022034816A (ja) * | 2020-08-19 | 2022-03-04 | ヤフー株式会社 | 情報処理装置、情報処理方法および情報処理プログラム |
| JP2023036804A (ja) * | 2021-03-05 | 2023-03-14 | 凸版印刷株式会社 | 電子チラシ管理装置、電子チラシ管理方法 |
| JP7292040B2 (ja) | 2019-01-17 | 2023-06-16 | ヤフー株式会社 | 情報処理プログラム、情報処理装置及び情報処理方法 |
Families Citing this family (17)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP6499477B2 (ja) * | 2015-02-27 | 2019-04-10 | 日本放送協会 | オントロジー生成装置、メタデータ出力装置、コンテンツ取得装置、オントロジー生成方法及びオントロジー生成プログラム |
| US10613521B2 (en) | 2016-06-09 | 2020-04-07 | Rockwell Automation Technologies, Inc. | Scalable analytics architecture for automation control systems |
| US10509396B2 (en) | 2016-06-09 | 2019-12-17 | Rockwell Automation Technologies, Inc. | Scalable analytics architecture for automation control systems |
| US10528700B2 (en) | 2017-04-17 | 2020-01-07 | Rockwell Automation Technologies, Inc. | Industrial automation information contextualization method and system |
| US10620612B2 (en) * | 2017-06-08 | 2020-04-14 | Rockwell Automation Technologies, Inc. | Predictive maintenance and process supervision using a scalable industrial analytics platform |
| US10990601B1 (en) * | 2018-03-12 | 2021-04-27 | A9.Com, Inc. | Dynamic optimization of variant recommendations |
| CN110362686B (zh) * | 2018-04-02 | 2024-02-06 | 北京搜狗科技发展有限公司 | 一种词库的生成方法、装置、终端设备和服务器 |
| US11144042B2 (en) | 2018-07-09 | 2021-10-12 | Rockwell Automation Technologies, Inc. | Industrial automation information contextualization method and system |
| US11403541B2 (en) | 2019-02-14 | 2022-08-02 | Rockwell Automation Technologies, Inc. | AI extensions and intelligent model validation for an industrial digital twin |
| US11086298B2 (en) | 2019-04-15 | 2021-08-10 | Rockwell Automation Technologies, Inc. | Smart gateway platform for industrial internet of things |
| US11435726B2 (en) | 2019-09-30 | 2022-09-06 | Rockwell Automation Technologies, Inc. | Contextualization of industrial data at the device level |
| US11841699B2 (en) | 2019-09-30 | 2023-12-12 | Rockwell Automation Technologies, Inc. | Artificial intelligence channel for industrial automation |
| US11249462B2 (en) | 2020-01-06 | 2022-02-15 | Rockwell Automation Technologies, Inc. | Industrial data services platform |
| US11501067B1 (en) * | 2020-04-23 | 2022-11-15 | Wells Fargo Bank, N.A. | Systems and methods for screening data instances based on a target text of a target corpus |
| US11726459B2 (en) | 2020-06-18 | 2023-08-15 | Rockwell Automation Technologies, Inc. | Industrial automation control program generation from computer-aided design |
| CN112015897B (zh) * | 2020-08-27 | 2023-04-07 | 中国平安人寿保险股份有限公司 | 语料的意图标注方法、装置、设备及存储介质 |
| CN115858781A (zh) * | 2022-11-29 | 2023-03-28 | 重庆长安汽车股份有限公司 | 一种文本标签提取方法、装置、设备及介质 |
Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH07325827A (ja) * | 1994-04-07 | 1995-12-12 | Mitsubishi Electric Corp | ハイパーテキスト自動生成装置 |
| JP2009026195A (ja) * | 2007-07-23 | 2009-02-05 | Yokohama National Univ | 商品分類装置、商品分類方法及びプログラム |
| JP2009181408A (ja) * | 2008-01-31 | 2009-08-13 | Nippon Telegr & Teleph Corp <Ntt> | 単語意味付与装置、単語意味付与方法、プログラムおよび記録媒体 |
Family Cites Families (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP4521343B2 (ja) | 2005-09-29 | 2010-08-11 | 株式会社東芝 | 文書処理装置及び文書処理方法 |
| US8438469B1 (en) * | 2005-09-30 | 2013-05-07 | Google Inc. | Embedded review and rating information |
| US8359191B2 (en) * | 2008-08-01 | 2013-01-22 | International Business Machines Corporation | Deriving ontology based on linguistics and community tag clouds |
-
2013
- 2013-09-30 JP JP2014508630A patent/JP5576003B1/ja active Active
- 2013-09-30 WO PCT/JP2013/076545 patent/WO2015045155A1/ja not_active Ceased
- 2013-09-30 US US14/420,424 patent/US9645979B2/en active Active
Patent Citations (3)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH07325827A (ja) * | 1994-04-07 | 1995-12-12 | Mitsubishi Electric Corp | ハイパーテキスト自動生成装置 |
| JP2009026195A (ja) * | 2007-07-23 | 2009-02-05 | Yokohama National Univ | 商品分類装置、商品分類方法及びプログラム |
| JP2009181408A (ja) * | 2008-01-31 | 2009-08-13 | Nippon Telegr & Teleph Corp <Ntt> | 単語意味付与装置、単語意味付与方法、プログラムおよび記録媒体 |
Cited By (5)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP7292040B2 (ja) | 2019-01-17 | 2023-06-16 | ヤフー株式会社 | 情報処理プログラム、情報処理装置及び情報処理方法 |
| JP2022034816A (ja) * | 2020-08-19 | 2022-03-04 | ヤフー株式会社 | 情報処理装置、情報処理方法および情報処理プログラム |
| JP7136856B2 (ja) | 2020-08-19 | 2022-09-13 | ヤフー株式会社 | 情報処理装置、情報処理方法および情報処理プログラム |
| JP2023036804A (ja) * | 2021-03-05 | 2023-03-14 | 凸版印刷株式会社 | 電子チラシ管理装置、電子チラシ管理方法 |
| JP7327628B2 (ja) | 2021-03-05 | 2023-08-16 | 凸版印刷株式会社 | 電子チラシ管理装置、電子チラシ管理方法 |
Also Published As
| Publication number | Publication date |
|---|---|
| US20160041951A1 (en) | 2016-02-11 |
| US9645979B2 (en) | 2017-05-09 |
| JPWO2015045155A1 (ja) | 2017-03-02 |
| JP5576003B1 (ja) | 2014-08-20 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| JP5576003B1 (ja) | コーパス生成装置、コーパス生成方法、及びコーパス生成プログラム | |
| JP6462970B1 (ja) | 分類装置、分類方法、生成方法、分類プログラム及び生成プログラム | |
| US8271513B1 (en) | Module and method for searching named entity of terms from the named entity database using named entity database and mining rule merged ontology schema | |
| JP6505421B2 (ja) | 情報抽出支援装置、方法およびプログラム | |
| CN113868568A (zh) | 一种网页关键字高亮方法、装置、设备及存储介质 | |
| US20090024616A1 (en) | Content retrieving device and retrieving method | |
| JP2002032770A (ja) | 文書処理方法、文書処理システムおよび媒体 | |
| CN103608805B (zh) | 辞典产生装置及方法 | |
| KR20150084706A (ko) | 온톨로지의 지식 학습 장치 및 그의 방법 | |
| WO2018070026A1 (ja) | 商品情報表示システム、商品情報表示方法、及びプログラム | |
| JP5324018B1 (ja) | コーパス生成装置、コーパス生成方法及びコーパス生成プログラム | |
| US11113314B2 (en) | Similarity calculating device and method, and recording medium | |
| JP2021009591A (ja) | データ取得装置、データ取得方法、およびデータ取得プログラム | |
| JP5085584B2 (ja) | 記事特徴語抽出装置、記事特徴語抽出方法及びプログラム | |
| JP2009205499A (ja) | ウェブページ特定装置、ウェブページ特定方法およびウェブページ特定用プログラム | |
| JP7795869B2 (ja) | 情報表現構造解析装置、および情報表現構造解析方法 | |
| JP5903171B2 (ja) | データ加工システムおよびデータ加工方法 | |
| JP2013143021A (ja) | 商品情報抽出ルール生成方法、装置、及びプログラム | |
| JP5739352B2 (ja) | 辞書生成装置、文書ラベル判定システム及びコンピュータプログラム | |
| JP5289468B2 (ja) | 回答検索装置、方法、及びプログラム | |
| JP5184987B2 (ja) | 索引情報作成装置、索引情報作成方法及びプログラム | |
| CN114661678B (zh) | 一种移除多语言资源的方法、装置、设备及介质 | |
| JPWO2016147624A1 (ja) | 検索システム、検索方法および検索プログラム | |
| JP2007200252A (ja) | 省略語生成・妥当性評価方法、同義語データベース生成・更新方法、省略語生成・妥当性評価装置、同義語データベース生成・更新装置、プログラム、記録媒体 | |
| JP5380130B2 (ja) | ファイル検索装置及びファイル検索方法、並びにプログラム |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| ENP | Entry into the national phase |
Ref document number: 2014508630 Country of ref document: JP Kind code of ref document: A |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 14420424 Country of ref document: US |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application |
Ref document number: 13894510 Country of ref document: EP Kind code of ref document: A1 |
|
| NENP | Non-entry into the national phase |
Ref country code: DE |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 13894510 Country of ref document: EP Kind code of ref document: A1 |