WO2014102992A1 - データ加工システムおよびデータ加工方法 - Google Patents

データ加工システムおよびデータ加工方法 Download PDF

Info

Publication number
WO2014102992A1
WO2014102992A1 PCT/JP2012/084007 JP2012084007W WO2014102992A1 WO 2014102992 A1 WO2014102992 A1 WO 2014102992A1 JP 2012084007 W JP2012084007 W JP 2012084007W WO 2014102992 A1 WO2014102992 A1 WO 2014102992A1
Authority
WO
WIPO (PCT)
Prior art keywords
information
meta information
meta
data
dictionary
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2012/084007
Other languages
English (en)
French (fr)
Inventor
藤田 雄介
信尾 額賀
児玉 昇司
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hitachi Ltd
Original Assignee
Hitachi Ltd
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hitachi Ltd filed Critical Hitachi Ltd
Priority to US14/649,762 priority Critical patent/US20150324436A1/en
Priority to JP2014553983A priority patent/JP5903171B2/ja
Priority to PCT/JP2012/084007 priority patent/WO2014102992A1/ja
Publication of WO2014102992A1 publication Critical patent/WO2014102992A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/25Integrating or interfacing systems involving database management systems
    • G06F16/254Extract, transform and load [ETL] procedures, e.g. ETL data flows in data warehouses
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/20Information retrieval; Database structures therefor; File system structures therefor of structured data, e.g. relational data
    • G06F16/28Databases characterised by their database models, e.g. relational or object models
    • G06F16/284Relational databases
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/90Details of database functions independent of the retrieved data types
    • G06F16/93Document management systems
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F7/00Methods or arrangements for processing data by operating upon the order or content of the data handled
    • G06F7/22Arrangements for sorting or merging computer data on continuous record carriers, e.g. tape, drum, disc
    • G06F7/24Sorting, i.e. extracting data from one or more carriers, rearranging the data in numerical or other ordered sequence, and rerecording the sorted data on the original carrier or on a different carrier or set of carriers sorting methods in general

Definitions

  • the present invention relates to data extraction and processing technology from a plurality of modalities.
  • Patent Document 1 includes "a method, system and apparatus including a computer program stored on a computer storage medium for retrieving and displaying information from collected electronic documents.
  • One aspect is to provide an existing structured presentation. By comparing the behavior of receiving descriptive data to describe and the characteristics of an existing structured presentation with the contents of the electronic document contained in the collected unstructured electronic document, new attributes related to the existing structured presentation can be obtained.
  • An action to identify the electronic document to be shown, an action to form an extended structured presentation by adding a new attribute identifier to an existing structured presentation, and an instruction to present the extended structured presentation Can be performed by a method implemented in a computer, including Summary reference).
  • Patent Document 1 US Patent Application Publication No. 2010/0185934
  • data mining for predicting the occurrence of an event using data is performed based on structural data arranged in a table or a relational database.
  • the data that can be imported as structural data is only data that has been previously given an attribute name and attribute value in a computer system for a specific application, and is not structured such as an image, sound, or an atypical document. Data cannot be directly targeted for data mining.
  • the full-text search engine for text documents can search for words in unstructured data at high speed, and simple conditional search using a word list is possible.
  • the above-mentioned Patent Document 1 discloses a method of combining these techniques and adding a table attribute name and attribute value.
  • the present invention has been made in view of these points, and an object thereof is to provide a data extraction / processing method for performing data mining / conditional search using data extracted from a plurality of modalities. There is to do.
  • the present invention provides a data processing system having one or more processors and one or more storage devices connected to the one or more processors.
  • the dictionary information for extracting meta information that defines the conditions for extracting meta information from the data of the data
  • the relevance dictionary information that defines the conditions for associating the meta information extracted from the plurality of types of data.
  • the meta information is extracted from the type of data based on the dictionary information for meta information extraction, the meta information is extracted from the input data, and is extracted from the input data based on the association dictionary information
  • the associated meta information and the meta information extracted from the plurality of types of data, and based on the result of the association, the plurality of types of data and the input data When the meta information extracted from these data, and outputs information indicating the association of any combination of.
  • meta information related to input data can be easily searched and processed.
  • Example 2 of this invention It is a figure explaining the detail of the process in which the text meta information extraction part extracts the meta information with respect to text data in Example 2 of this invention. It is a flowchart which shows the operation
  • an example of a data processing system that extends an input table based on a meta information database relating to image data and document data that has been constructed in advance will be described.
  • the system can be used for managing design drawings and design documents issued when a building or a machine is manufactured, for example.
  • a table related to design is input, the table is automatically expanded using meta information extracted from the design drawing and the design document. Since a large-scale table related to design can be obtained in this way, this embodiment can be applied to data mining such as design failure analysis or failure prediction.
  • FIG. 1 is a block diagram illustrating the overall configuration of the data processing system according to the first embodiment of the present invention.
  • the data processing system 1 includes a data source server 2, an ETL (Extract Transform Load) server 3, a storage server 4, a meta information extraction server 5, a meta information search server 6, and a data processing server 7.
  • ETL Extract Transform Load
  • the data source server 2 is a device that manages images and documents.
  • the data source server 2 includes a relational database (not shown) that manages drawings in association with IDs (identification information), and a file server (not shown) that stores text documents.
  • the ETL server 3 has a function of storing image data and document data stored in the data source server 2 in the storage server 4. Here, conversion such as unifying the format of the image and the document is performed.
  • the storage server 4 includes an image data storage unit 11 and a document data storage unit 12, and stores image data and document data collected from a plurality of data sources in a unified format.
  • the meta information extraction server 5 includes an image dictionary unit 13, an image meta information extraction unit 14, a document dictionary unit 15, a document meta information extraction unit 16, a relationship dictionary unit 17, a meta information association unit 18, and a meta information database 19. And manages meta information extracted from data in the storage server 4.
  • the meta information search server 6 includes a related meta information search unit 20, receives a search request, and returns a result of searching the meta information database 19.
  • the data processing server 7 includes a data input unit 21, an input data association unit 22, a data processing unit 23, and a data output unit 24, and processes and outputs input data based on meta information.
  • FIG. 2 is a block diagram illustrating a hardware configuration of the data processing system 1 according to the first embodiment of the present invention.
  • the data source server 2 is a computer having a communication unit 221, a CPU (Central Processing Unit) 222, a memory 223, and a disk 224 that are connected to each other.
  • a communication unit 221 a CPU (Central Processing Unit) 222, a memory 223, and a disk 224 that are connected to each other.
  • CPU Central Processing Unit
  • the communication unit 221 is an interface that is connected to the relay device 280 and communicates with other servers via the relay device 280.
  • the CPU 222 is a processor that realizes a predetermined function by executing a program stored in the memory 223.
  • the memory 223 and the disk 224 are storage devices that store programs executed by the CPU 222, data referred to by the CPU 222, and the like. These may be any type of storage device.
  • the memory 223 is a relatively high-speed semiconductor memory such as a DRAM (Dynamic Random Access Memory), and the disk 224 is a hard disk. It is a relatively large capacity storage device such as a device.
  • the ETL server 3 is a computer having a communication unit 231, a CPU 232, a memory 233, and a disk 234 connected to each other.
  • the description of these units is the same as the description of the communication unit 221, the CPU 222, the memory 223, and the disk 224 of the data source server 2, and will be omitted.
  • the storage server 4 is a computer having a communication unit 241, a CPU 242, a memory 243, and a disk 244 connected to each other.
  • the description of these units is the same as the description of the communication unit 221, the CPU 222, the memory 223, and the disk 224 of the data source server 2, and will be omitted.
  • the disk 244 includes an image data storage unit 11 that stores image data and a document data storage unit 12 that stores document data. Part or all of the data stored in the image data storage unit 11 and the document data storage unit 12 may be copied to the memory 243 as necessary.
  • the meta information extraction server 5 is a computer having a communication unit 251, a CPU 252, a memory 253, and a disk 254 that are connected to each other.
  • the description of these units is the same as the description of the communication unit 221, the CPU 222, the memory 223, and the disk 224 of the data source server 2, and will be omitted.
  • the memory 253 includes an image meta information extraction unit 14, a document meta information extraction unit 16, and a meta information association unit 18. These are programs executed by the CPU 252.
  • processing executed by the image meta information extraction unit 14, the document meta information extraction unit 16, or the meta information association unit 18 may be described. However, such processing is actually required by the CPU 252 according to the above program. Accordingly, the process is executed by controlling the memory 253, the disk 254, the communication unit 251, and the like. Processing executed by programs stored in a memory 263 and a memory 273, which will be described later, is actually executed by the CPU of each computer in the same manner as described above.
  • the image meta information extraction unit 14, the document meta information extraction unit 16, and the meta information association unit 18 may be stored in the disk 254 and copied to the memory 253 as necessary. The same applies to programs stored in a memory 263 and a memory 273 described later.
  • the meta information search server 6 is a computer having a communication unit 261, a CPU 262, a memory 263, and a disk 264 that are connected to each other.
  • the description of these units is the same as the description of the communication unit 221, the CPU 222, the memory 223, and the disk 224 of the data source server 2, and will be omitted.
  • the memory 263 includes the related meta information search unit 20. This is a program executed by the CPU 262.
  • the data processing server 7 is a computer having a communication unit 271, a CPU 272, a memory 273, and a disk 274 that are connected to each other.
  • the description of these units is the same as the description of the communication unit 221, the CPU 222, the memory 223, and the disk 224 of the data source server 2, and will be omitted.
  • the memory 273 includes a data input unit 21, an input data association unit 22, a data processing unit 23, and a data output unit 24. These are programs executed by the CPU 272.
  • the data processing server 7 further includes an input unit 275 and an output unit 276 that are connected to the CPU 272 and controlled by the data input unit 21 and the data output unit 24.
  • the input unit 275 is an input device such as a keyboard and a pointing device
  • the output unit 276 is an output device such as an image display device.
  • each server other than the data processing server 7 constituting the data processing system 1 may have the same input unit and output unit as the data processing server 7.
  • the relay device 280 is a device that is connected to the communication unit of each server and relays communication between servers.
  • FIG. 2 shows an example of a hardware configuration realized by an independent computer in which each server includes one CPU and one or more storage devices.
  • each server includes one CPU and one or more storage devices.
  • the data processing system 1 may be realized by one computer including a storage device that stores all the programs and data described above and at least one CPU.
  • the data source server 2 is realized by one computer
  • the ETL server 3 and the storage server 4 are realized by another one computer
  • the meta information extraction server 5, the meta information search server 6, and the data processing server 7 are further provided. It may be realized by another single computer.
  • each server may be a virtual server generated using a virtualization technique.
  • FIG. 3 is a flowchart showing the operation of the meta information database construction process in Embodiment 1 of the present invention.
  • the ETL server 3 acquires image data and document data from the data source server 2 (step S301). Subsequently, the ETL server 3 performs necessary conversion on the image data and the document data (step S302). For example, when the later-described image meta information extracting unit 14 accepts only a specific format of image data, a process of converting the format of the image data into the specific format is performed. Subsequently, the ETL server 3 stores the converted image data and document data in the image data storage unit 11 and the document data storage unit 12 of the storage server 4 (step S303).
  • the meta information extraction server 5 acquires image data and document data from the storage server 4 (step S304).
  • the image meta information extraction unit 14 of the meta information extraction server 5 extracts meta information for the image data based on the image dictionary unit 13 (step S305).
  • FIG. 4 is a diagram illustrating details of the process (step S305 in FIG. 3) in which the image meta information extraction unit 14 extracts meta information for image data in the first embodiment of the present invention.
  • the image dictionary unit 13 holds a model 401 for performing recognition and classification based on image shape / color information and the like.
  • Each model 401 is associated with a label 402.
  • “shape: tube” is associated as a label 402 with a model 401 corresponding to an image of a tube-shaped figure.
  • the image meta information extraction unit 14 gives a label 402 to the image data 403 acquired from the image data storage unit 11 by an image recognition technique based on the image dictionary unit 13. Specifically, for example, the image meta information extraction unit 14 compares the similarity between the image data 403 and each model 401 by a known image recognition technique, and the label 402 corresponding to the model 401 having the highest similarity is displayed as an image. You may give to the data 403. In the example of FIG. 4, the image meta information extraction unit 14 assigns a label of “shape: tube” to the image data 403 and outputs the result as image meta information 404.
  • the image meta information 404 is an expression indicating that “shape of C01.jpg is a tube”. For example, an expression of a ternary relationship such as “B of A is C” using RDF (Resource Description Framework). Can be described.
  • step S 305 the document meta information extraction unit 16 extracts meta information for the document data based on the document dictionary unit 15.
  • FIG. 5 is a diagram for explaining the details of the process (step S305) in which the document meta information extraction unit 16 extracts meta information for the document data in the first embodiment of the present invention.
  • the document dictionary unit 15 holds an attribute name list 501 in which words suitable for attribute names are collected and an attribute value list 502 in which words suitable for attribute values are collected for the words in the document.
  • the document meta information extraction unit 16 analyzes the structure of the document data 503 acquired from the document data storage unit 12 based on the document dictionary unit 15 and the result of analyzing the layout of the document, thereby obtaining the document meta information 504. Is generated.
  • the document data 503 includes character strings such as “notification: C01” and “location: Tokyo”.
  • the attribute name list 501 includes words such as “reference”, “location”, and “creator”, and the attribute value list 502 includes words such as “C01”, “C02”, “Tokyo”, and “Alice”.
  • the document meta information extraction unit 16 extracts attribute names such as “reference” and “location” from the document data 503 based on the attribute name list 501, and “C01” and “ By extracting attribute values such as “Tokyo” and associating, for example, “reference” with “C01”, “location” with “Tokyo” based on the layout information that the words are arranged with a colon in between, As a result of the structuring, document meta information 504 is generated.
  • the document meta information 504 can be described using RDF, similarly to the image meta information 404.
  • the image meta information extraction unit 14 and the document meta information extraction unit 16 store the extracted meta information in the meta information database 19 (step S306).
  • image meta information and document meta information can be managed in the same database.
  • the meta information associating unit 18 uses the relevance dictionary unit 17 to give association to the meta information stored in the meta information database 19 (step S307).
  • FIG. 6 is a diagram for explaining the details of the process (step S307) in which the meta information association unit 18 assigns the association to the meta information in the first embodiment of the present invention.
  • the relevancy dictionary unit 17 holds a synonym dictionary 601, and examples of synonym relationships include “drawing” and “reference”, “type” and “shape”, “C01” and “C01.jpg”, and the like. Pre-built by the system designer.
  • the synonym dictionary 601 may hold a synonym relationship between words in different languages based on a translation dictionary prepared separately.
  • the synonym dictionary may be any information as long as it is information defining information that can be replaced with certain meta information.
  • the synonym dictionary is information indicating that information included in meta information of a modal and information included in meta information of another modal that is synonymous with the information can be replaced.
  • a dictionary (see Example 2) that defines a synonym relationship between spoken words and written words may be used.
  • the meta information associating unit 18 searches the meta information database 19 for a word existing in the synonym dictionary 601 and converts the word into changed meta information 602 in which the word is replaced with a word having a synonym relation. Update.
  • the image meta information “shape of C01.jpg is tube” is searched for the word “shape” in the synonym dictionary 601.
  • the meta information associating unit 18 replaces “shape” with “type” that is the synonym based on the synonym dictionary 601, whereby the searched image meta information is “type of C01.jpg is tube. Is converted to meta information.
  • step S307 for example, according to the synonym relation, the word “shape” and the word “type” in the meta information database can be replaced, and the word “C01” and the word “C01.jpg” can be replaced.
  • the meta information associating unit 18 replaces the searched meta information “shape of C01.jpg is tube” with the word “shape” and the word “shape” based on the synonym dictionary 601.
  • Information indicating that the “type” can be replaced and information indicating that the word “C01” and the word “C01.jpg” can be replaced may be added to the meta information.
  • the image meta information and the document meta information are associated with each other via “C01.jpg”. That is, “_: r1 drawing is C01.jpg and C01.jpg type is tube”, in other words, “_: r1 drawing type is tube”.
  • “drawing-type” obtained by combining two attribute names “drawing” and “type” associated via the attribute value “C01.jpg” may be treated as a new attribute name (see FIG. 10).
  • the meta information database 19 is constructed as described above.
  • meta information extracted from multiple modal data contains the same (or related) concept, but different expressions are given to each modal meta information depending on the nature of the modal
  • These expressions are unified by updating the meta information as described above or adding meta information that can be replaced. This makes it possible to search related meta information without omission in the processing described later.
  • FIG. 7 is a flowchart showing an operation of data processing in the first embodiment of the present invention.
  • the data input unit 21 receives input of table data from the user (step S701).
  • the table data input here (for example, input data 801 to be described later) is stored in advance in the data processing system 1 (or acquired from the outside of the data processing system 1 via, for example, the input unit 275 or the communication unit 271).
  • Arbitrary structural data which is associated with unstructured data by the following process.
  • the input data association unit 22 extracts attribute relationships from the received table (step S702).
  • FIG. 8 is a diagram for explaining the details of the process (step S702) for extracting the attribute relationship from the table received by the input data association unit 22 in the first embodiment of the present invention.
  • the input data association unit 22 extracts record information from the input data 801, extracts the relationship between attributes and attribute names, and outputs table structure attribute information 802.
  • the table structure attribute information 802 can be described using RDF, similarly to the description of image meta information and document meta information.
  • “M001”, “A”, “JP1”, and “C01” are included in one record of the input data 801 as values corresponding to the items “ID”, “part”, “location”, and “drawing”, respectively. If it is included, the record is identified as “_: r2”, “_: r2 ID is M001”, “_: r2 component is A”, “_: r2 location is JP1 Table structure attribute information 802 including RDF description such as “A certain” and “_: r2 drawing is C01” is output.
  • the input data associating unit 22 assigns an association to the table structure attribute information 802 using the relevancy dictionary unit 17 (step S703).
  • FIG. 9 is a diagram illustrating details of the process (step S703) in which the input data association unit 22 assigns association to the table structure attribute information 802 in the first embodiment of the present invention.
  • the input data association unit 22 searches the table structure attribute information 802 for a word existing in the synonym dictionary 601 and replaces the word with a word having a synonym relationship with the word.
  • Change attribute information 901 is generated, and the table structure attribute information 802 is changed based on the change attribute information 901.
  • “_: r2 drawing is C01” is searched from the table structure attribute information 802 for the word “C01” in the synonym dictionary 601, and these are “C01” and “C01.
  • the attribute information is changed to “_: r2 drawing is C01.jpg”.
  • the input data association unit 22 can replace “C01” and “C01.jpg” instead of rewriting the table structure attribute information 802 with the change attribute information 901. May be added to the table structure attribute information 802.
  • the related meta information search unit 20 searches the meta information database 19 using the table structure attribute information 802 (step S704).
  • meta information having the same attribute relationship as the input is selected as an inquiry to the meta information database 19.
  • the meta information database 19 includes information “_: r1 drawing is C01.jpg” (( It can be found as the same attribute relationship as the input. Similarly, “the location of _: r2 is Tokyo” can also be found for “the location of _: r2 is Tokyo”.
  • the related meta-information search unit 20 can regard them as the same record when the attribute relationship of a plurality of records satisfies a predetermined condition. For example, if a plurality of records have two or more common attribute relationships (that is, a combination of a common attribute name and attribute value) and are regarded as the same record, in the above example, _: r1 _: R2 can be regarded as the same record.
  • the condition for regarding two records as the same is not limited to the above example.
  • a list of attribute names regarded as the same record is defined in advance, and the related meta information search unit 20 regards a plurality of records having the same attribute value corresponding to the attribute name included in the list as the same record. Good.
  • the related meta information search unit 20 may estimate the same record from the search result of meta information.
  • the related meta-information search unit 20 acquires the attribute relationship regarding _: r1, which is regarded as the same record as _: r2 included in the input data, from the meta-information database 19 (step S705).
  • _: r1 which is regarded as the same record as _: r2 included in the input data
  • the meta-information database 19 the meta-information database 19
  • two pieces of meta information “the creator of _: r1 is Alice” and “the delivery date of _: r1 is 2012/7/30” can be acquired (see FIG. 6).
  • the related meta-information search unit 20 further traces any attribute related to the attribute value, thereby saying that “_: r1 drawing (is C01.jpg and C01.jpg) is a tube”. Meta information can also be acquired.
  • the related meta-information search unit 20 “_: r2 creator is Alice”, “_: r2 delivery date is 2012/7/30”, and “_: r2 drawing (C01. jpg and C01.jpg) is a tube "is determined as additional attribute information.
  • the data processing unit 23 adds an attribute to the input data 801 based on the additional attribute information determined by the related meta information search unit 20 (step S706).
  • FIG. 10 is a diagram illustrating processing (step S706) in which the data processing unit 23 adds an attribute to the input data 801 in the first embodiment of the present invention.
  • the data processing unit 23 adds the additional attribute information 1002 to the table structure attribute information 802 and returns it to the table representation, whereby the processed data 1003 is obtained.
  • “_: r2 creator is Alice
  • “_: r2 delivery date is 2012/7/30”
  • “_: r2 drawing (is C01.jpg, "C01.jpg) type is tube” is added in tabular form.
  • another attribute name is associated with an attribute value corresponding to a certain attribute name as in the relationship between the above “drawing”, “C01.jpg”, and “type”
  • the plurality of attribute names is connected and displayed as a new attribute name.
  • Such processed data 1003 indicates the relationship between the input data 801, the image meta information 404, and the document meta information 504.
  • the data output unit 24 outputs machining data 1003.
  • the input table is extended based on a meta information database that has been constructed in advance.
  • by expressing the meta information in RDF it is possible to associate data of different modalities such as a table, a document, and an image based on a single synonym dictionary.
  • the meta information about the image type is used, but the result of other information extraction can be used.
  • the extracted person name may be used as meta information using face image recognition.
  • an example of a data processing system that presents information related to input text data based on a meta information database extracted from voice data and text data will be described.
  • the system can be used, for example, to manage call recording voices and operator response logs stored in a call center.
  • the system displays the related voice and the response log by using the recorded voice and the meta information extracted from the response log. This makes it possible to search for information efficiently without listening to all recorded audio.
  • FIG. 11 is a block diagram illustrating the overall configuration of the data processing system according to the second embodiment of the present invention.
  • the image data and document data in the first embodiment are replaced with audio data and text data, and a meta information database is constructed for each of the audio data and the text data. Instead of associating, the meta information is searched when it is associated.
  • each part of the data processing system of the second embodiment has the same functions as the parts denoted by the same reference numerals of the first embodiment shown in FIG. Description of is omitted.
  • the data processing system 1 includes the data source server 2, the ETL server 3, the storage server 4, the meta information extraction server 5, the meta information search server 6, and the data processing server 7 as in the first exemplary embodiment. .
  • the data source server 2 is a device that manages voice and text, and includes a relational database that manages recorded voice data in association with an ID, and a file server that stores voice data files and text files.
  • the ETL server 3 has a function of storing voice data and text data stored in the data source server 2 in the storage server 4. Here, conversion such as unifying the format of the audio data is performed.
  • the storage server 4 includes an audio data storage unit 51 and a text data storage unit 52, and stores audio data and text data collected from a plurality of data sources in a unified format.
  • the meta information extraction server 5 includes an audio dictionary unit 53, an audio meta information extraction unit 54, a text dictionary unit 55, a text meta information extraction unit 56, an audio meta information database 57, and a text meta information database 58.
  • the meta information extracted from the data in is managed.
  • the meta information search server 6 includes a relevancy dictionary unit 17, a meta information association unit 18, and a related meta information search unit 20.
  • the meta information search server 6 receives the search request and searches the speech meta information database 57 and the text meta information database 58, respectively. return it.
  • the data processing server 7 includes a data input unit 21, an input data association unit 22, a data processing unit 23, and a data output unit 24, and processes and outputs the input data based on the meta information.
  • the hardware configuration of the data processing system 1 of the present embodiment is the same as the configuration of the first embodiment shown in FIG.
  • the disk 244 of the storage server 4 includes an audio data storage unit 51 and a text data storage unit 52 instead of the image data storage unit 11 and the document data storage unit 12.
  • the disk 254 of the meta information extraction server 5 is replaced with an audio dictionary unit 53, a text dictionary unit 55, an audio meta, instead of the image dictionary unit 13, the document dictionary unit 15, the relevance dictionary unit 17, and the meta information database 19.
  • An information database 57 and a text meta information database 58 are included.
  • the memory 253 of the meta information extraction server 5 includes an audio meta information extraction unit 54 and a text meta information extraction unit 56 instead of the image meta information extraction unit 14, the document meta information extraction unit 16, and the meta information association unit 18.
  • the memory 263 of the meta information search server 6 includes the meta information association unit 18 in addition to the related meta information search unit 20, and the disk 264 includes the relevancy dictionary unit 17.
  • the operation of the data processing system 1 is divided into a meta information database construction process and a data processing process.
  • the meta information database construction process is the same as the meta information database construction process in the first embodiment shown in FIG. 3 except for the differences described below.
  • the ETL server 3 acquires voice data and text data from the data source server 2 (step S301). Subsequently, the ETL server 3 performs necessary conversion on the voice data and the text data (step S302). Subsequently, the ETL server 3 stores the converted voice data and text data in the storage server 4 (step S303).
  • the meta information extraction server 5 acquires voice data and text data from the storage server 4 (step S304).
  • the audio meta information extraction unit 54 extracts meta information for the audio data based on the audio dictionary unit 53 (step S305).
  • FIG. 12 is a diagram illustrating details of the process (step S305) in which the audio meta information extraction unit 54 extracts meta information for audio data in the second embodiment of the present invention.
  • the voice dictionary unit 53 holds a keyword list 1201 that is a list of keywords detected from the voice.
  • the speech meta information extraction unit 54 generates speech meta information 1204 by dividing the speech data 1203 into sentence units by using a speech recognition technique based on the speech dictionary unit 53 and adding keywords appearing therein.
  • two keywords “product A” and “will go” are extracted from the audio data 1203.
  • the audio meta-information 1204 includes “W01.wav sentence has _: s1”, “W01.wav sentence has _: s2”, and “_: s1 keyword has“ product A ”. "There is” and "_: s2 keyword has” I will "will be described.
  • step S305 the text meta information extraction unit 56 extracts meta information for the text data based on the text dictionary unit 55.
  • FIG. 13 is a diagram for explaining the details of the process (step S305) in which the text meta information extraction unit 56 extracts meta information for text data in the second embodiment of the present invention.
  • the text dictionary unit 55 holds a keyword list 1301 extracted from the text.
  • the text meta information extraction unit 56 generates text meta information 1304 by analyzing the text data 1303 based on the text dictionary unit 55 and the result of the morphological analysis. In the example of FIG. 13, three keywords “product A”, “claim”, and “visit” are extracted from the text data 1303.
  • the text meta information 1304 can be described using RDF, like the audio meta information 1204.
  • the audio meta information extraction unit 54 stores the extracted audio meta information in the audio meta information database 57 (step S306).
  • the text meta information extraction unit 56 stores the extracted text meta information in the text meta information database 58.
  • the extracted meta information is stored as a database separated by each modal.
  • step S307 for associating meta information is not performed after storing meta information (step S306).
  • FIG. 14 is a flowchart showing the data processing operation in the second embodiment of the present invention.
  • the data input unit 21 receives input of text data from the user (step S1401).
  • the input data association unit 22 extracts keywords from the received text data (step S1402).
  • a case where “machine ABC visit” is input as text data and the keywords “machine ABC” and “visit” are extracted will be described as an example.
  • the input data association unit 22 may extract the keyword using morphological analysis or the like.
  • the input data association unit 22 uses the association dictionary unit 17 to associate the extracted keywords (step S1403).
  • FIG. 15 is a diagram for explaining the details of the process (step S1403) for assigning the keyword extracted by the input data association unit 22 in the second embodiment of the present invention.
  • a synonym dictionary 601 including information for converting between spoken words and written words is built in advance. This makes it possible to associate “visit” with “will go”.
  • the synonym dictionary 601 illustrated in FIG. 15 further includes information that associates “machine ABC” with “product A”.
  • “Machine ABC” and “Product A” are aliases of the same product (for example, one is identification information used inside the manufacturer and the other is a product name used for the customer). It is.
  • the input data association unit 22 uses the keywords “machine ABC” and “visit” extracted from the input text data 1501 as “product A” and “ Can be associated with.
  • the meta information associating unit 18 develops the extracted keywords based on the result of the association using the relevancy dictionary unit 17 (step S1404). Specifically, the meta information association unit 18 uses the extracted keywords “machine ABC” and “visit” as keywords for searching the audio meta information database 57 (that is, a search query for audio) “product A” and “ Will expand. Similarly, the meta information association unit 18 uses the extracted keywords “machine ABC” and “visit” as keywords for searching the text meta information database 58 (that is, a text search query) “product A” and “visit”. Expand to.
  • the related meta-information search unit 20 searches the speech meta-information database 57 using the keywords “product A” and “will go”, and further searches the text meta-information database 58 using the keywords “product A” and “visit”. (Step S1405).
  • the data processing unit 23 associates the audio meta information searched by the related meta information search unit 20 with the text meta information (step S1406).
  • FIG. 16 is a diagram illustrating processing (step S1406) in which the data processing unit 23 associates audio meta information and text meta information in the second embodiment of the present invention.
  • step 1401 when the text data 1501 “machine ABC visit” is input in step 1401 and the audio meta information 1204 and the text meta information 1304 are searched in step S 1405, the data processing unit 23 stores the text data corresponding to the text meta information 1304.
  • Processing data 1601 is generated by associating 1303 with audio data 1203 corresponding to the audio meta information 1204.
  • the data processing unit 23 applies the keywords “product A” and “visit” included in the audio data 1203 to the keywords “product A” and “visit” included in the text data 1303, respectively.
  • Processing data 1601 is generated by adding links (“listen W01.wav@s1” and “listen W01.wav@s2” in the example of FIG. 16), respectively.
  • the processed data 1601 is information indicating a relationship between a keyword corresponding to the input keyword included in the text data 1303 and a portion of the audio data 1203 corresponding to the keyword.
  • the data output unit 24 outputs the processed data 1601 (step S1407).
  • the speech meta information database 57 and the text meta information database 58 are not associated in advance. Therefore, these databases may contain information on the same or related concepts, given different expressions depending on the modal (for example, expressions such as spoken and written).
  • the input keyword is expanded based on the relevance dictionary unit 17 to convert the input keyword into an expression that matches each modal feature, whether each modal feature (for example, spoken or written) is included. ) Can be searched for.
  • processing such as generating a link of related voice data on the text of the call answering record.
  • a dictionary that defines a combination of replaceable words such as a synonym dictionary or a spoken word conversion dictionary, as a relevance dictionary that is the same as the word association in the meta information database and the word association in the input table.
  • Information extracted from multimodal data such as data, image data, and audio data can be associated with the input data.
  • various functions of the data processing system are realized by a program executed on the CPU of each server.
  • some or all of them include electronic components such as integrated circuits. It may be realized by the hardware used.
  • the present invention is not limited to the above-described embodiment, and includes various modifications.
  • a table generation system for data mining and a voice log analysis system for call centers are assumed.
  • a system for managing electronic medical record data and medical image data, a broadcast data editing system, etc. It can be applied to various systems for managing modal data.
  • Information such as a program, a table, and a file that realize each function of the above embodiment is a storage device such as a nonvolatile semiconductor memory, a hard disk drive, an SSD (Solid State Drive), or an IC card, an SD card, a DVD, or the like. It can be stored on a computer readable non-transitory data storage medium.
  • a storage device such as a nonvolatile semiconductor memory, a hard disk drive, an SSD (Solid State Drive), or an IC card, an SD card, a DVD, or the like. It can be stored on a computer readable non-transitory data storage medium.
  • the present invention is not limited to the above-described embodiment, and includes various modifications.
  • the above-described embodiment has been described in detail for easy understanding of the present invention, and is not necessarily limited to one having all the configurations described.
  • a part of the configuration of an embodiment can be replaced with the configuration of another embodiment, and the configuration of another embodiment can be added to the configuration of an embodiment.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Databases & Information Systems (AREA)
  • General Engineering & Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Data Mining & Analysis (AREA)
  • Business, Economics & Management (AREA)
  • General Business, Economics & Management (AREA)
  • Computer Hardware Design (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

 一つ以上のプロセッサと、前記一つ以上のプロセッサに接続される一つ以上の記憶装置と、を有するデータ加工システムであって、複数の種類のデータからメタ情報を抽出する条件を定義するメタ情報抽出用辞書情報と、前記複数の種類のデータから抽出されたメタ情報を関連付ける条件を定義する関連性辞書情報と、を保持し、前記複数の種類のデータから、前記メタ情報抽出用辞書情報に基づいて前記メタ情報を抽出し、入力されたデータからメタ情報を抽出し、前記関連性辞書情報に基づいて、前記入力されたデータから抽出されたメタ情報と前記複数の種類のデータから抽出されたメタ情報とを関連付け、前記関連付けの結果に基づいて、前記複数の種類のデータと、前記入力されたデータと、それらのデータから抽出されたメタ情報と、のいずれかの組み合わせの関連を示す情報を出力する。

Description

データ加工システムおよびデータ加工方法
 本発明は、複数のモダリティからのデータ抽出および加工技術に関する。
 本技術分野の背景技術として、米国特許出願公開第2010/0185934号明細書(特許文献1)がある。この明細書には、「収集された電子文書から情報を検索及び表示するための、計算機記憶媒体に格納された計算機プログラムを含む方法、システム及び装置。一つの態様は、既存の構造化プレゼンテーションを記述する記述データを受信する動作と、既存の構造化プレゼンテーションの特性と収集された非構造化電子文書に含まれる電子文書の内容とを比較することで、既存の構造化プレゼンテーションに関する新たな属性を示す電子文書を特定する動作と、新たな属性の識別子を既存の構造化プレゼンテーションに追加することで、拡張された構造化プレゼンテーションを形成する動作と、拡張された構造化プレゼンテーションを提示する指示を出力する動作と、を含む、計算機に実装された方法によって実施され得る」と記載されている(要約参照)。
 特許文献1:米国特許出願公開第2010/0185934号明細書
 従来、データを用いて事象の発生を予測するためのデータマイニングは、表または関係データベースといった形に整理された構造データに基づいて行われる。しかし、構造データとして取り込むことのできるデータは、特定用途向けに、コンピュータシステム内で予め属性名と属性値の付与が行われたものだけであって、画像、音声または非定型の文書といった非構造データを、直接データマイニングの対象とすることはできない。
 一方、テキスト文書を対象とした全文検索エンジンは、非構造データ中の単語を高速に検索することができ、単語リストを用いての単純な条件付き検索は可能となった。また、複数の構造データをルールに基づいて結び付けることは可能であり、インターネット上の大量のデータから構造データを検索して、大きな構造データを生成することが可能となった。さらに、非構造データに対し、テキストの構文解析を行うことで、部分的な構造を取得することは可能となった。例えば、上記の特許文献1は、これらの技術を組み合わせ、表の属性名および属性値を追加する方法を開示している。
 しかしながら、非構造データを構造データに結び付けたデータマイニング、および単語リストだけでない、非構造データに対する属性条件を使った検索は、実現されていない。また、非構造データに対して、部分的に構造を与えることができても、構造データ中のどの行または列に結び付けるかを判断する手法は従来なかった。
 本発明は、このような点に鑑みてなされたものであり、その目的は、複数のモダリティから抽出したデータを用いて、データマイニング・条件付き検索を行うための、データ抽出・加工方法を提供することにある。
 上記の課題を解決するために、本発明は、一つ以上のプロセッサと、前記一つ以上のプロセッサに接続される一つ以上の記憶装置と、を有するデータ加工システムであって、複数の種類のデータからメタ情報を抽出する条件を定義するメタ情報抽出用辞書情報と、前記複数の種類のデータから抽出されたメタ情報を関連付ける条件を定義する関連性辞書情報と、を保持し、前記複数の種類のデータから、前記メタ情報抽出用辞書情報に基づいて前記メタ情報を抽出し、入力されたデータからメタ情報を抽出し、前記関連性辞書情報に基づいて、前記入力されたデータから抽出されたメタ情報と前記複数の種類のデータから抽出されたメタ情報とを関連付け、前記関連付けの結果に基づいて、前記複数の種類のデータと、前記入力されたデータと、それらのデータから抽出されたメタ情報と、のいずれかの組み合わせの関連を示す情報を出力する。
 本発明の一実施形態によれば、入力データに関連するメタ情報を容易に検索して、加工することができる。
 上記以外の課題、構成及び効果は、以下の実施形態の説明により明らかにされる。
本発明の実施例1のデータ加工システムの全体構成を説明するブロック図である。 本発明の実施例1のデータ加工システムのハードウェア構成を説明するブロック図である。 本発明の実施例1におけるメタ情報データベース構築処理の動作を示すフローチャートである。 本発明の実施例1において画像メタ情報抽出部が画像データに対するメタ情報を抽出する処理の詳細を説明する図である。 本発明の実施例1において文書メタ情報抽出部が文書データに対するメタ情報を抽出する処理の詳細を説明する図である。 本発明の実施例1においてメタ情報関連付け部がメタ情報に関連付けを付与する処理の詳細を説明する図である。 本発明の実施例1におけるデータ加工処理の動作を示すフローチャートである。 本発明の実施例1において入力データ関連付け部が受け付けた表から属性の関係を抽出する処理の詳細を説明する図である。 本発明の実施例1において入力データ関連付け部が表構造属性情報に関連付けを付与する処理の詳細を説明する図である。 本発明の実施例1においてデータ加工部が入力データに属性を追加する処理を説明する図である。 本発明の実施例2のデータ加工システムの全体構成を説明するブロック図である。 本発明の実施例2において音声メタ情報抽出部が音声データに対するメタ情報を抽出する処理の詳細を説明する図である。 本発明の実施例2においてテキストメタ情報抽出部がテキストデータに対するメタ情報を抽出する処理の詳細を説明する図である。 本発明の実施例2におけるデータ加工処理の動作を示すフローチャートである。 本発明の実施例2において入力データ関連付け部が抽出されたキーワードに関連付けを付与する処理の詳細を説明する図である。 本発明の実施例2においてデータ加工部が音声メタ情報とテキストメタ情報を関連付ける処理を説明する図である。
 以下、実施例を、図面を用いて説明する。
 本実施例では、予め構築しておいた画像データおよび文書データに関するメタ情報データベースに基づいて、入力された表を拡張する、データ加工システムの例を説明する。本システムは、例えば、建造物または機械等を製造する際に発行される、設計図面および設計文書を管理するために用いることができる。設計に関する表を入力すると、設計図面および設計文書から抽出されたメタ情報を用いて、表が自動的に拡張される。こうして設計に関わる大規模な表が得られるため、本実施例は、設計の不具合分析または不良予測など、データマイニングに応用することが可能となる。
 図1は、本発明の実施例1のデータ加工システムの全体構成を説明するブロック図である。
 データ加工システム1は、データソースサーバ2、ETL(Extract Transform Load)サーバ3、ストレージサーバ4、メタ情報抽出サーバ5、メタ情報検索サーバ6、およびデータ加工サーバ7によって構成される。
 データソースサーバ2は、画像および文書を管理する装置である。データソースサーバ2は、図面をID(識別情報)と結び付けて管理するリレーショナルデータベース(図示省略)、およびテキスト文書を保存するファイルサーバ(図示省略)を備える。
 ETLサーバ3は、データソースサーバ2に保存されている画像データおよび文書データを、ストレージサーバ4に保存する機能を備える。ここで、画像および文書のフォーマットを統一するなどの変換が行われる。
 ストレージサーバ4は、画像データ保存部11および文書データ保存部12を備え、複数のデータソースから収集された画像データおよび文書データを統一された形式で保存する。
 メタ情報抽出サーバ5は、画像用辞書部13、画像メタ情報抽出部14、文書用辞書部15、文書メタ情報抽出部16、関連性辞書部17、メタ情報関連付け部18、およびメタ情報データベース19を備え、ストレージサーバ4にあるデータから抽出したメタ情報の管理を行う。
 メタ情報検索サーバ6は、関連メタ情報検索部20を備え、検索要求を受けつけてメタ情報データベース19を検索した結果を返す。
 データ加工サーバ7は、データ入力部21、入力データ関連付け部22、データ加工部23、およびデータ出力部24を備え、入力されたデータをメタ情報に基づいて加工し出力する。
 上記の各部の詳細については後述する。
 図2は、本発明の実施例1のデータ加工システム1のハードウェア構成を説明するブロック図である。
 データソースサーバ2は、相互に接続された通信部221、CPU(Central Processing Unit)222、メモリ223およびディスク224を有する計算機である。
 通信部221は、中継装置280に接続され、中継装置280を介して他のサーバと通信するためのインターフェースである。CPU222は、メモリ223に格納されたプログラムを実行することによって所定の機能を実現するプロセッサである。メモリ223及びディスク224は、CPU222によって実行されるプログラムおよびCPU222によって参照されるデータ等を格納する記憶装置である。これらはどのような種類の記憶装置であってもよいが、典型的な例を示すと、メモリ223がDRAM(Dynamic Random Access Memory)のような比較的高速の半導体メモリであり、ディスク224がハードディスク装置のような比較的大容量の記憶装置である。
 ETLサーバ3は、相互に接続された通信部231、CPU232、メモリ233およびディスク234を有する計算機である。これらの各部の説明は、データソースサーバ2の通信部221、CPU222、メモリ223およびディスク224の説明と同様であるため、省略する。
 ストレージサーバ4は、相互に接続された通信部241、CPU242、メモリ243およびディスク244を有する計算機である。これらの各部の説明は、データソースサーバ2の通信部221、CPU222、メモリ223およびディスク224の説明と同様であるため、省略する。ただし、ディスク244は、画像データを保存する画像データ保存部11および文書データを保存する文書データ保存部12を含む。画像データ保存部11および文書データ保存部12に保存されたデータの一部又は全部が必要に応じてメモリ243にコピーされてもよい。
 メタ情報抽出サーバ5は、相互に接続された通信部251、CPU252、メモリ253およびディスク254を有する計算機である。これらの各部の説明は、データソースサーバ2の通信部221、CPU222、メモリ223およびディスク224の説明と同様であるため、省略する。ただし、メモリ253は、画像メタ情報抽出部14、文書メタ情報抽出部16およびメタ情報関連付け部18を含む。これらは、CPU252によって実行されるプログラムである。
 以下、画像メタ情報抽出部14、文書メタ情報抽出部16またはメタ情報関連付け部18が実行する処理について説明する場合があるが、そのような処理は、実際には上記のプログラムに従ってCPU252が必要に応じてメモリ253、ディスク254及び通信部251等を制御することによって実行する処理である。後述するメモリ263及びメモリ273に格納されたプログラムが実行する処理も、実際には、上記と同様にそれぞれの計算機のCPUによって実行される。
 なお、画像メタ情報抽出部14、文書メタ情報抽出部16およびメタ情報関連付け部18は、ディスク254に格納され、必要に応じてメモリ253にコピーされてもよい。後述するメモリ263及びメモリ273に格納されたプログラムについても同様である。
 メタ情報検索サーバ6は、相互に接続された通信部261、CPU262、メモリ263およびディスク264を有する計算機である。これらの各部の説明は、データソースサーバ2の通信部221、CPU222、メモリ223およびディスク224の説明と同様であるため、省略する。ただし、メモリ263は、関連メタ情報検索部20を含む。これは、CPU262によって実行されるプログラムである。
 データ加工サーバ7は、相互に接続された通信部271、CPU272、メモリ273およびディスク274を有する計算機である。これらの各部の説明は、データソースサーバ2の通信部221、CPU222、メモリ223およびディスク224の説明と同様であるため、省略する。ただし、メモリ273は、データ入力部21、入力データ関連付け部22、データ加工部23及びデータ出力部24を含む。これらは、CPU272によって実行されるプログラムである。
 データ加工サーバ7は、さらに、CPU272に接続され、データ入力部21およびデータ出力部24によって制御される入力部275および出力部276を有する。入力部275は、例えばキーボードおよびポインティングデバイスのような入力デバイスであり、出力部276は、例えば画像表示装置のような出力デバイスである。
 図2では省略されているが、データ加工システム1を構成するデータ加工サーバ7以外の各サーバも、データ加工サーバ7と同様の入力部および出力部を有してもよい。
 中継装置280は、各サーバの通信部に接続され、サーバ間の通信を中継する装置である。
 図2には、各サーバが一つのCPUおよび一つ以上の記憶装置を備える独立した計算機によって実現されるハードウェア構成の例を示したが、このようなハードウェア構成は一例であり、実際には一つ以上のCPUおよび一つ以上の記憶装置を有する種々の形態の計算機システムによって本実施例を実現することができる。例えば、上記の全てのプログラム及びデータを格納する記憶装置と、少なくとも一つのCPUとを含む一つの計算機によってデータ加工システム1が実現されてもよい。あるいは、例えばデータソースサーバ2が一つの計算機によって実現され、ETLサーバ3及びストレージサーバ4が別の一つの計算機によって実現され、メタ情報抽出サーバ5、メタ情報検索サーバ6及びデータ加工サーバ7がさらに別の一つの計算機によって実現されてもよい。このような場合、各サーバは、仮想化技術を利用して生成された仮想サーバであってもよい。
 次に、上記のように構成される、本実施例に係るデータ加工システム1の動作を説明する。本システムの動作は、メタ情報データベース構築処理とデータ加工処理とに分けられる。
 まず、メタ情報データベースの構築処理に関する動作を説明する。
 図3は、本発明の実施例1におけるメタ情報データベース構築処理の動作を示すフローチャートである。
 まず、ETLサーバ3は、データソースサーバ2から、画像データおよび文書データを取得する(ステップS301)。続いて、ETLサーバ3は、画像データおよび文書データに対し、必要な変換を施す(ステップS302)。例えば、後述する画像メタ情報抽出部14が、特定の形式の画像データしか受け付けない場合には、画像データの形式を当該特定の形式に変換する処理が行われる。続いて、ETLサーバ3は、変換された画像データおよび文書データを、それぞれ、ストレージサーバ4の画像データ保存部11及び文書データ保存部12へ保存する(ステップS303)。
 次に、メタ情報抽出サーバ5は、ストレージサーバ4から画像データおよび文書データを取得する(ステップS304)。
 次に、メタ情報抽出サーバ5の画像メタ情報抽出部14は、画像用辞書部13に基づいて、画像データに対するメタ情報を抽出する(ステップS305)。
 図4は、本発明の実施例1において画像メタ情報抽出部14が画像データに対するメタ情報を抽出する処理(図3のステップS305)の詳細を説明する図である。
 画像用辞書部13は、画像の形状・色情報等に基づいて認識・分類を行うためのモデル401を保持している。また、各モデル401には、ラベル402が対応づけられている。例えば筒(tube)状の図形の画像に相当するモデル401にはラベル402として“shape:tube”が対応づけられる。画像メタ情報抽出部14は、画像用辞書部13に基づく画像認識技術によって、画像データ保存部11から取得された画像データ403に対するラベル402を付与する。具体的には、例えば、画像メタ情報抽出部14は、画像データ403と各モデル401との類似度を公知の画像認識技術によって比較し、最も類似度の高いモデル401に対応するラベル402を画像データ403に付与してもよい。図4の例では、画像メタ情報抽出部14は、画像データ403に“shape:tube”のラベルを付与し、その結果を、画像メタ情報404として出力する。
 画像メタ情報404は、「C01.jpgのshapeがtubeである」ことを示す表現であり、例えば、RDF(Resource Description Framework)を用いて「AのBはCである」といった三項関係の表現を記述することができる。
 さらに、ステップS305において、文書メタ情報抽出部16は、文書用辞書部15に基づいて、文書データに対するメタ情報を抽出する。
 図5は、本発明の実施例1において文書メタ情報抽出部16が文書データに対するメタ情報を抽出する処理(ステップS305)の詳細を説明する図である。
 文書用辞書部15は、文書中の単語に対し、属性名にふさわしい単語を集めた属性名リスト501および属性値にふさわしい単語を集めた属性値リスト502を保持している。文書メタ情報抽出部16は、文書用辞書部15と文書のレイアウトを解析した結果とに基づいて、文書データ保存部12から取得された文書データ503の構造を解析することによって、文書メタ情報504を生成する。
 図5の例では、文書データ503は、「通知:C01」および「場所:東京」等の文字列を含む。一方、属性名リスト501は「参照」、「場所」および「作成者」等の単語を含み、属性値リスト502は「C01」、「C02」、「東京」および「Alice」等の単語を含む。
 この例において、文書メタ情報抽出部16は、文書データ503から、属性名リスト501に基づいて「参照」および「場所」といった属性名を抽出し、属性値リスト502に基づいて「C01」および「東京」といった属性値を抽出し、それらの単語がコロンを挟んで並べられているといったレイアウト情報に基づいて例えば「参照」と「C01」、「場所」と「東京」を対応付けることによって、表の構造化を行った結果として文書メタ情報504を生成する。文書メタ情報504は、画像メタ情報404と同様に、RDFを用いて記述することができる。
 次に、画像メタ情報抽出部14および文書メタ情報抽出部16は、抽出したメタ情報を、メタ情報データベース19に保存する(ステップS306)。ここで、RDFを管理するデータベースを用いると、画像メタ情報と文書メタ情報を同一のデータベースの中で管理することができる。
 次に、メタ情報関連付け部18は、関連性辞書部17を用いて、メタ情報データベース19の中に保存されているメタ情報に関連付けを付与する(ステップS307)。
 図6は、本発明の実施例1においてメタ情報関連付け部18がメタ情報に関連付けを付与する処理(ステップS307)の詳細を説明する図である。
 関連性辞書部17は、同義語辞書601を保持しており、同義語関係として、例えば「図面」と「参照」、「タイプ」と「shape」、「C01」と「C01.jpg」などが、システムの設計者によって予め構築される。同義語辞書601は、別途用意された翻訳辞書に基づいて、異なる言語の単語間の同義関係を保持するものであってもよい。
 なお、同義語辞書は、あるメタ情報と置換可能な情報を定義する情報である限り、どのようなものであってもよい。具体的には、同義語辞書は、あるモーダルのメタ情報に含まれる情報と、それと同義の、他のモーダルのメタ情報に含まれる情報とが置換可能であることを示す情報であり、例えば、上記のような同義語辞書または翻訳辞書のほか、話し言葉と書き言葉との同義関係を定義する辞書(実施例2参照)であってもよい。
 メタ情報関連付け部18は、この同義語辞書601に存在する単語をメタ情報データベース19から検索し、その単語を同義語関係にある単語に置き換えた変更メタ情報602に変換し、メタ情報データベース19を更新する。図6の例では、同義語辞書601内の単語「shape」に対して、「C01.jpgのshapeはtubeである」という画像メタ情報が検索される。この場合、メタ情報関連付け部18は、同義語辞書601に基づいて「shape」をその同義語である「タイプ」に置き換えることによって、検索された画像メタ情報を「C01.jpgのタイプはtubeである」というメタ情報に変換する。
 同様にして、「_:r1の参照はC01である」が「_:r1の図面はC01.jpgである」に、「_:r1の場所は東京である」が「_:r1の場所はJP1である」に変換される。
 ここでは、検索されたメタ情報を新たなメタ情報に更新する例を説明したが、これは、同義語関係に従って、検索されたメタ情報に含まれる属性名または属性値の単語と置換可能な単語を記述する方法の一例である。ステップS307では、例えば、同義語関係に従って、メタ情報データベース中の単語「shape」と単語「タイプ」が置換可能であることと、単語「C01」と単語「C01.jpg」が置換可能であることを、記述することができればよい。具体的には、メタ情報関連付け部18は、例えば、検索されたメタ情報「C01.jpgのshapeはtubeである」を書き換える代わりに、同義語辞書601に基づいて、単語「shape」と単語「タイプ」が置換可能であることを示す情報、及び、単語「C01」と単語「C01.jpg」が置換可能であることを示す情報を当該メタ情報に追加してもよい。
 なお、上記のように検索されたメタ情報を同義語辞書に基づいて新たなメタ情報に更新する場合には、同義関係にあるメタ情報に単一の表現が与えられるように更新する必要がある。例えば、あるメタ情報に単語「C01」が含まれ、別のメタ情報に単語「C01.jpg」が含まれ、さらに別のメタ情報にそれらと同義の別の単語が含まれる場合に、単語「C01.jpg」と同義の単語が全てそれらを代表する単語「C01.jpg」に変換される必要がある。どの単語が代表として使用されてもよく、例えば同義関係にある複数の単語のリストの先頭の単語が使用されてもよい。
 ここで、図6中に点線で示したように、「C01.jpg」を介して画像メタ情報と文書メタ情報とが関連付けられる。すなわち、「_:r1の図面はC01.jpgであり、C01.jpgのタイプはtubeである」言い換えると「_:r1の図面のタイプはtubeである」という意味付けがされたことになる。このような場合、属性値「C01.jpg」を介して関連付けられた二つの属性名「図面」および「タイプ」を結合した「図面-タイプ」が新たな属性名として扱われてもよい(図10参照)。
 以上のようにして、メタ情報データベース19が構築される。
 例えば、複数のモーダルのデータから抽出されたメタ情報が同一の(または関連する)概念を含んでいるが、モーダルの性質に応じてそれぞれのモーダルのメタ情報に異なる表現が与えられている場合に、上記のようなメタ情報の更新又は置換可能なメタ情報の追加等によって、それらの表現が統一される。これによって、後述する処理において、関連するメタ情報を漏れなく検索することが可能になる。
 次に、データ加工処理に関する動作を説明する。
 図7は、本発明の実施例1におけるデータ加工処理の動作を示すフローチャートである。
 まず、データ入力部21は、ユーザから表データの入力を受けつける(ステップS701)。ここで入力される表データ(例えば後述する入力データ801)は、データ加工システム1の内部にあらかじめ保持された(または、例えば入力部275もしくは通信部271を介してデータ加工システム1の外部から取得された)任意の構造データであり、これが以下の処理によって非構造データと対応付けられる。
 次に、入力データ関連付け部22は、受け付けた表から属性の関係を抽出する(ステップS702)。
 図8は、本発明の実施例1において入力データ関連付け部22が受け付けた表から属性の関係を抽出する処理(ステップS702)の詳細を説明する図である。
 入力データ関連付け部は22、入力データ801に対し、レコードの情報を抽出し、属性と属性名の関係を抽出し、表構造属性情報802を出力する。表構造属性情報802は、画像メタ情報および文書メタ情報の記述と同様に、RDFを用いて記述することができる。
 例えば、入力データ801の一つのレコードに、項目「ID」、「部品」、「場所」および「図面」のそれぞれに対応する値として「M001」、「A」、「JP1」および「C01」が含まれる場合、そのレコードを「_:r2」と識別して、「_:r2のIDはM001である」、「_:r2の部品はAである」、「_:r2の場所はJP1である」および「_:r2の図面はC01である」といったRDFによる記述を含む表構造属性情報802が出力される。
 次に、入力データ関連付け部22は、関連性辞書部17を用いて、表構造属性情報802に関連付けを付与する(ステップS703)。
 図9は、本発明の実施例1において入力データ関連付け部22が表構造属性情報802に関連付けを付与する処理(ステップS703)の詳細を説明する図である。
 メタ情報関連付け部18の機能と同様、入力データ関連付け部22は、同義語辞書601に存在する単語を表構造属性情報802から検索し、その単語をそれと同義語関係にある単語に置き変えることによって変更属性情報901を生成し、変更属性情報901に基づいて表構造属性情報802を変更する。図9の例では、同義語辞書601内の単語「C01」に対して、表構造属性情報802から「_:r2の図面はC01である」が検索され、それが「C01」と「C01.jpg」との同義語関係に基づいて「_:r2の図面がC01.jpgである」という属性情報に変更される。
 なお、メタ情報関連付け部18が実行する関連付け処理と同様に、入力データ関連付け部22は、表構造属性情報802を変更属性情報901に書き換える代わりに、「C01」と「C01.jpg」が置換可能であることを示す情報を表構造属性情報802に追加してもよい。
 次に、関連メタ情報検索部20は、表構造属性情報802を用いて、メタ情報データベース19を検索する(ステップS704)。ここでは、メタ情報データベース19に対する問い合わせとして、入力と同じ属性関係を持つメタ情報が選択される。表構造属性情報802に存在する「_:r2の図面はC01.jpgである」に対して、メタ情報データベース19に「_:r1の図面はC01.jpgである」という情報が含まれるため(図6参照)、それを入力と同一の属性関係として見つけることができる。また同様に、「_:r2の場所は東京である」に対して「_:r1の場所は東京である」も見つけることができる。
 ここで、関連メタ情報検索部20は、複数のレコードの属性関係が所定の条件を満たす場合に、それらを同一のレコードと見なすことができる。例えば、複数のレコードが二つ以上の共通する属性関係(すなわち共通する属性名と属性値との組)を有する場合にそれらを同一のレコードと見なすとすれば、上記の例では_:r1と_:r2とを同一レコードと見なせる。なお、二つのレコードを同一と見なすための条件は、上記の例に限定されない。例えば、同一レコードとみなす属性名のリストを予め定義し、関連メタ情報検索部20は、そのリストに含まれる属性名に対応する属性値が同一である複数のレコードを同一のレコードと見なしてもよい。あるいは、関連メタ情報検索部20は、メタ情報の検索結果から同一のレコードを推定してもよい。
 次に、関連メタ情報検索部20は、入力データに含まれる_:r2と同一レコードとみなされた_:r1に関する属性関係をメタ情報データベース19から取得する(ステップS705)。ここでは、「_:r1の作成者はAliceである」、「_:r1の納期は2012/7/30である」の2つのメタ情報が取得できる(図6参照)。さらに、関連メタ情報検索部20は、属性値に関する属性があればそれをさらにたどることによって、「_:r1の図面(はC01.jpgであり、C01.jpg)のタイプはtubeである」というメタ情報も取得できる。こうして、関連メタ情報検索部20は、「_:r2の作成者はAliceである」、「_:r2の納期は2012/7/30である」、および「_:r2の図面(はC01.jpgであり、C01.jpg)のタイプはtubeである」の3つの属性情報を追加属性情報として決定する。
 次に、データ加工部23は、関連メタ情報検索部20により決定された追加属性情報に基づいて、入力データ801に属性を追加する(ステップS706)。
 図10は、本発明の実施例1においてデータ加工部23が入力データ801に属性を追加する処理(ステップS706)を説明する図である。
 データ加工部23が表構造属性情報802に対して追加属性情報1002を追加し、表の表現に戻すことによって、加工データ1003が得られる。図10の例では、「_:r2の作成者はAliceである」、「_:r2の納期は2012/7/30である」、および「_:r2の図面(はC01.jpgであり、C01.jpg)のタイプはtubeである」の3つの属性情報が表形式で追加される。ここで、上記の「図面」と「C01.jpg」と「タイプ」との関係のように、ある属性名に対応する属性値にさらに別の属性名が関連付けられる場合、それらの複数の属性名表現を繋げて新たな属性名として表示する。図10の例では、「図面」の「タイプ」という複数の属性名表現をハイフン記号で繋げることによって、「図面‐タイプ」というように表現される。このような加工データ1003によって、入力データ801と、画像メタ情報404と、文書メタ情報504との関連が示される。
 次に、データ出力部24は、加工データ1003を出力する。
 以上のようにして、データ加工処理が行われる。
 本実施例では、予め構築しておいたメタ情報データベースに基づいて、入力された表を拡張する一例を示した。メタ情報データベースにおける単語の関連付けと、入力された表内の単語の関連付けに、同一の関連性辞書として同義語辞書を用いることによって、文書データや画像データから抽出された情報を、入力された表に関連付けることが可能となる。また、メタ情報をRDFで表現しておくことによって、表と文書と画像という異なるモダリティのデータを、単一の同義語辞書に基づいて関連付けることが可能となる。
 また、本実施例では、画像のタイプについてのメタ情報を利用したが、他の情報抽出による結果を用いることもできる。例えば顔写真のような画像に対しては、顔画像認識を用いて、抽出された人名をメタ情報として利用してもよい。
 以下、本発明の実施例2を、図面を用いて説明する。
 本実施例では、音声データとテキストデータから抽出されたメタ情報データベースに基づいて、入力されたテキストデータに関連する情報を提示する、データ加工システムの例を説明する。本システムは、例えば、コールセンターにおいて蓄積される、通話録音音声とオペレータの応対ログを管理するために用いることができる。本システムは、テキストを入力されると、録音音声と応対ログから抽出されたメタ情報を用いて、関連する音声および応対ログを表示する。これによって、全ての録音音声を聴くことなく、効率的に情報を探索することが可能となる。
 図11は、本発明の実施例2のデータ加工システムの全体構成を説明するブロック図である。
 本実施例は、実施例1における画像データと文書データを、音声データとテキストデータに置き換え、メタ情報データベースを、音声データ用とテキストデータ用でそれぞれ構築しておき、メタ情報データベース構築時にこれらの関連付けを行なわず、メタ情報検索時に関連付ける構成とした。
 上記およびこれから詳細に説明する相違点を除き、実施例2のデータ加工システムの各部は、図1に示された実施例1の同一の符号を付された各部と同一の機能を有するため、それらの説明は省略する。
 本実施例のデータ加工システム1は、実施例1と同様に、データソースサーバ2、ETLサーバ3、ストレージサーバ4、メタ情報抽出サーバ5、メタ情報検索サーバ6およびデータ加工サーバ7によって構成される。
 データソースサーバ2は、音声およびテキストを管理する装置であり、録音音声データをIDと結び付けて管理するリレーショナルデータベースと、音声データファイルとテキストファイルを保存するファイルサーバとを備える。
 ETLサーバ3は、データソースサーバ2に保存されている音声データおよびテキストデータを、ストレージサーバ4に保存する機能を備える。ここで、音声データのフォーマットを統一するなどの変換が行われる。
 ストレージサーバ4は、音声データ保存部51およびテキストデータ保存部52を備え、複数のデータソースから収集された音声データおよびテキストデータを統一された形式で保存する。
 メタ情報抽出サーバ5は、音声用辞書部53、音声メタ情報抽出部54、テキスト用辞書部55、テキストメタ情報抽出部56、音声メタ情報データベース57およびテキストメタ情報データベース58を備え、ストレージサーバ4にあるデータから抽出したメタ情報の管理を行う。
 メタ情報検索サーバ6は、関連性辞書部17、メタ情報関連付け部18および関連メタ情報検索部20を備え、検索要求を受けつけて、音声メタ情報データベース57およびテキストメタ情報データベース58をそれぞれ検索した結果を返す。
 データ加工サーバ7は、データ入力部21、入力データ関連付け部22、データ加工部23およびデータ出力部24を備え、入力されたデータをメタ情報に基づいて加工し出力する。
 本実施例のデータ加工システム1のハードウェア構成は、図2に示した実施例1の構成と同様である。ただし、ストレージサーバ4のディスク244は、画像データ保存部11および文書データ保存部12の代わりに音声データ保存部51およびテキストデータ保存部52を含む。メタ情報抽出サーバ5のディスク254は、画像用辞書部13、文書用辞書部15、関連性辞書部17およびメタ情報データベース19の代わりに、音声用辞書部53、テキスト用辞書部55、音声メタ情報データベース57およびテキストメタ情報データベース58を含む。メタ情報抽出サーバ5のメモリ253は、画像メタ情報抽出部14、文書メタ情報抽出部16およびメタ情報関連付け部18の代わりに、音声メタ情報抽出部54およびテキストメタ情報抽出部56を含む。メタ情報検索サーバ6のメモリ263は、関連メタ情報検索部20に加えてメタ情報関連付け部18を含み、ディスク264は関連性辞書部17を含む。
 次に、上記のように構成される、本実施例に係るデータ加工システム1の動作を説明する。本システムの動作は、実施例1と同様に、メタ情報データベース構築処理とデータ加工処理に分けられる。メタ情報データベースの構築処理は、以下に説明する相違点を除き、図3に示した実施例1におけるメタ情報データベースの構築処理と同様である。
 まず、ETLサーバ3は、データソースサーバ2から、音声データとテキストデータを取得する(ステップS301)。続いて、ETLサーバ3は、音声データとテキストデータに対し、必要な変換を施す(ステップS302)。続いて、ETLサーバ3は、変換された音声データとテキストデータを、ストレージサーバ4へ保存する(ステップS303)。
 次に、メタ情報抽出サーバ5は、ストレージサーバ4から音声データとテキストデータを取得する(ステップS304)。
 次に、音声メタ情報抽出部54は、音声用辞書部53に基づいて、音声データに対するメタ情報を抽出する(ステップS305)。
 図12は、本発明の実施例2において音声メタ情報抽出部54が音声データに対するメタ情報を抽出する処理(ステップS305)の詳細を説明する図である。
 音声用辞書部53は、音声中から検出するキーワードのリストであるキーワードリスト1201を保持している。音声メタ情報抽出部54は、音声用辞書部53に基づく音声認識技術によって、音声データ1203を文単位に分割し、その中に出現するキーワードを付与することによって、音声メタ情報1204を生成する。図12の例では、音声データ1203から、2つのキーワード「製品A」と「まいります」が抽出される。音声メタ情報1204は、「W01.wavの文(sentence)には_:s1がある」、「W01.wavの文には_:s2がある」、「_:s1のキーワードには「製品A」がある」、「_:s2のキーワードには「まいります」がある」といった4つの関係を記述している。これらの関係は、実施例1と同様にRDFを用いて記述することができる。
 同様に、ステップS305において、テキストメタ情報抽出部56は、テキスト用辞書部55に基づいて、テキストデータに対するメタ情報を抽出する。
 図13は、本発明の実施例2においてテキストメタ情報抽出部56がテキストデータに対するメタ情報を抽出する処理(ステップS305)の詳細を説明する図である。
 テキスト用辞書部55は、テキスト中から抽出するキーワードリスト1301を保持している。テキストメタ情報抽出部56は、テキスト用辞書部55と、形態素解析の結果に基づいて、テキストデータ1303を解析することによって、テキストメタ情報1304を生成する。図13の例では、テキストデータ1303に対して、3つのキーワード「製品A」、「クレーム」および「訪問」が抽出される。テキストメタ情報1304は、音声メタ情報1204と同様に、RDFを用いて記述することができる。
 次に、音声メタ情報抽出部54は、抽出した音声メタ情報を音声メタ情報データベース57に保存する(ステップS306)。同様に、ステップS306において、テキストメタ情報抽出部56は、抽出したテキストメタ情報をテキストメタ情報データベース58に保存する。本実施例では、実施例1と異なり、抽出されたメタ情報が各モーダルで分離したデータベースとして保存される。
 本実施例では、実施例1と異なり、メタ情報保存(ステップS306)の後に、メタ情報関連付けを行うステップS307は行われない。
 次に、データ加工処理に関する動作を説明する。
 図14は、本発明の実施例2におけるデータ加工処理の動作を示すフローチャートである。
 まず、データ入力部21は、ユーザからテキストデータの入力を受けつける(ステップS1401)。
 次に、入力データ関連付け部22は、受け付けたテキストデータからキーワードを抽出する(ステップS1402)。ここでは、テキストデータとして「マシンABC 訪問」が入力され、キーワード「マシンABC」および「訪問」が抽出された場合を例として説明する。あるいは、例えばキーワードを含む自然文のテキストデータが入力された場合等に、入力データ関連付け部22は、形態素解析等を用いてキーワードを抽出してもよい。
 次に、入力データ関連付け部22は、関連性辞書部17を用いて、抽出されたキーワードに関連付けを付与する(ステップS1403)。
 図15は、本発明の実施例2において入力データ関連付け部22が抽出されたキーワードに関連付けを付与する処理(ステップS1403)の詳細を説明する図である。
 関連性辞書部17には、話し言葉と書き言葉を相互に変換するための情報等を含む同義語辞書601が予め構築されている。これによって「訪問」と「まいります」を対応付けることができる。図15に示す同義語辞書601は、さらに、「マシンABC」と「製品A」とを対応付ける情報を含む。この例において「マシンABC」および「製品A」は同一の製品の別名(例えば、一方が製造元の社内で使用される識別情報であり、もう一方が顧客向けに使用される商品名であるなど)である。
 入力データ関連付け部22は、関連性辞書部17に含まれる同義語辞書601に基づいて、入力されたテキストデータ1501から抽出されたキーワード「マシンABC」および「訪問」をそれぞれ「製品A」および「まいります」に関連付けることができる。
 次に、メタ情報関連付け部18は、上記の関連性辞書部17を用いた関連付けの結果に基づいて、抽出されたキーワードを展開する(ステップS1404)。具体的には、メタ情報関連付け部18は、抽出されたキーワード「マシンABC」および「訪問」を、音声メタ情報データベース57を検索するためのキーワード(すなわち音声用検索クエリ)「製品A」および「まいります」に展開する。同様に、メタ情報関連付け部18は、抽出されたキーワード「マシンABC」および「訪問」を、テキストメタ情報データベース58を検索するためのキーワード(すなわちテキスト用検索クエリ)「製品A」および「訪問」に展開する。
 次に、関連メタ情報検索部20は、キーワード「製品A」および「まいります」によって音声メタ情報データベース57を検索し、さらに、キーワード「製品A」および「訪問」によってテキストメタ情報データベース58を検索する(ステップS1405)。
 次に、データ加工部23は、関連メタ情報検索部20によって検索された音声メタ情報とテキストメタ情報を関連付ける(ステップS1406)。
 図16は、本発明の実施例2においてデータ加工部23が音声メタ情報とテキストメタ情報を関連付ける処理(ステップS1406)を説明する図である。
 例えば、ステップ1401においてテキストデータ1501「マシンABC 訪問」が入力され、ステップS1405において音声メタ情報1204およびテキストメタ情報1304が検索された場合、データ加工部23は、テキストメタ情報1304に対応するテキストデータ1303に、音声メタ情報1204に対応する音声データ1203を関連付けることによって加工データ1601を生成する。
 具体的には、例えば、データ加工部23は、テキストデータ1303に含まれるキーワード「製品A」および「訪問」に、それぞれ、音声データ1203に含まれるキーワード「製品A」および「まいります」へのリンク(図16の例ではそれぞれ「listen W01.wav@s1」および「listen W01.wav@s2」)を追加することによって、加工データ1601を生成する。言い換えると、加工データ1601は、テキストデータ1303に含まれる、入力されたキーワードに対応するキーワードと、それに対応する音声データ1203の部分と、の関連を示す情報である。
 次に、データ出力部24は、加工データ1601を出力する(ステップS1407)。
 以上のようにして、データ加工処理が行われる。
 本実施例では、実施例1と異なり、事前に音声メタ情報データベース57とテキストメタ情報データベース58との間の関連付けを行っていない。したがって、これらのデータベースは、モーダルに応じて異なる表現(例えば話し言葉と書き言葉等の表現)が与えられた、同一のまたは関連する概念の情報を含んでいる可能性がある。しかし、入力キーワードを関連性辞書部17に基づいて展開することで、入力キーワードが各モーダルの特徴に合わせた表現に変換されるため、各モーダルの特徴(例えば話し言葉または書き言葉のいずれが含まれるか)に合わせた検索が可能となる。その結果、例えば通話応対記録のテキスト上に、関連する音声データのリンクを生成する等の加工を行うことが可能となる。
 以上説明した本発明の各実施形態によれば、メタ情報データベースに基づいて、入力されたデータを加工することが可能となる。メタ情報データベースにおける単語の関連付けと、入力された表内の単語の関連付けに同一の関連性辞書として、同義語辞書または話し言葉変換辞書といった置換可能な単語の組み合わせを定義する辞書を用いることによって、文書データ・画像データ・音声データなど、マルチモーダルデータから抽出された情報を、入力されたデータに関連付けることが可能となる。
 なお、上述した各実施形態では、各サーバのCPU上で実行されるプログラムによって、データ加工システムの各種機能を実現しているが、それらの一部又は全部が、例えば集積回路等の電子部品を用いたハードウェアによって実現されてもよい。
 本発明は上述した実施形態に限定されるものではなく、様々な変形例が含まれる。本実施例では、データマイニングのための表生成システムや、コールセンターの音声ログ解析のシステムを想定したが、例えば、電子カルテデータと医用画像データを管理するシステムや、放送データの編集システムなど、マルチモーダルのデータを管理するための様々なシステムに適用することができる。
 上記の実施形態の各機能を実現するプログラム、テーブル、ファイル等の情報は、不揮発性半導体メモリ、ハードディスクドライブ、SSD(Solid State Drive)等の記憶デバイス、または、ICカード、SDカード、DVD等の計算機読み取り可能な非一時的データ記憶媒体に格納することができる。
 本発明は上記した実施形態に限定されるものではなく、様々な変形例が含まれる。例えば、上記した実施形態は本発明を分かりやすく説明するために詳細に説明したものであり、必ずしも説明した全ての構成を備えるものに限定されるものではない。また、ある実施形態の構成の一部を他の実施形態の構成に置き換えることが可能であり、また、ある実施形態の構成に他の実施形態の構成を加えることも可能である。また、各実施形態の構成の一部について、他の構成の追加・削除・置換をすることが可能である。

Claims (13)

  1.  一つ以上のプロセッサと、前記一つ以上のプロセッサに接続される一つ以上の記憶装置と、を有するデータ加工システムであって、
     複数の種類のデータからメタ情報を抽出する条件を定義するメタ情報抽出用辞書情報と、前記複数の種類のデータから抽出されたメタ情報を関連付ける条件を定義する関連性辞書情報と、を保持し、
     前記複数の種類のデータから、前記メタ情報抽出用辞書情報に基づいて前記メタ情報を抽出し、
     入力されたデータからメタ情報を抽出し、
     前記関連性辞書情報に基づいて、前記入力されたデータから抽出されたメタ情報と前記複数の種類のデータから抽出されたメタ情報とを関連付け、
     前記関連付けの結果に基づいて、前記複数の種類のデータと、前記入力されたデータと、それらのデータから抽出されたメタ情報と、のいずれかの組み合わせの関連を示す情報を出力することを特徴とするデータ加工システム。
  2.  請求項1に記載のデータ加工システムであって、
     前記複数の種類のデータは、第1の種類のデータおよび第2の種類のデータを含み、
     前記メタ情報抽出用辞書情報は、前記第1の種類のデータから第1のメタ情報を抽出する条件を定義する第1のメタ情報抽出用辞書情報と、前記第2の種類のデータから第1のメタ情報を抽出する条件を定義する第2のメタ情報抽出用辞書情報と、を含み、
     前記関連性辞書情報は、メタ情報と置換可能な情報を定義する情報を含み、
     前記データ加工システムは、
     前記第1の種類のデータおよび前記第2の種類のデータから、それぞれ、前記第1のメタ情報抽出用辞書情報および前記第2のメタ情報抽出用辞書情報に基づいて、前記第1のメタ情報および前記第2のメタ情報を抽出し、
     入力されたデータから第3のメタ情報を抽出し、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報、前記第2のメタ情報と置換可能な情報、および前記第3のメタ情報と置換可能な情報の少なくとも一つを特定し、
     前記第3のメタ情報または前記第3のメタ情報と置換可能な情報を用いて、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報、および、前記第2のメタ情報または前記第2のメタ情報と置換可能な情報を検索し、
     前記検索の結果に基づいて、前記第3のメタ情報または前記第3のメタ情報と置換可能な情報と、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報と、前記第2のメタ情報または前記第2のメタ情報と置換可能な情報と、を関連付けることを特徴とするデータ加工システム。
  3.  請求項2に記載のデータ加工システムであって、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報を特定し、前記関連性辞書情報に基づいて、前記第2のメタ情報と置換可能な情報を特定し、前記特定された置換可能な情報と前記第1のメタ情報と前記第2のメタ情報とが所定の条件を満たすか否かの判定に基づいて、前記第1のメタ情報と前記第2のメタ情報とを関連付け、
     前記関連性辞書情報に基づいて、前記第3のメタ情報と置換可能な情報を特定し、
     前記第1のメタ情報または前記第1のメタ情報と置換可能な情報が検索結果として取得された場合、前記第3のメタ情報または前記第3のメタ情報と置換可能な情報と、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報と、前記第1のメタ情報に関連付けられた前記第2のメタ情報または前記第2のメタ情報と置換可能な情報と、を関連付けることを特徴とするデータ加工システム。
  4.  請求項2に記載のデータ加工システムであって、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報を特定する手順、および、前記関連性辞書情報に基づいて、前記第2のメタ情報と置換可能な情報を特定する手順を実行せず、前記関連性辞書情報に基づいて、前記第3のメタ情報と置換可能な情報を特定することを特徴とするデータ加工システム。
  5.  請求項2に記載のデータ加工システムであって、
     前記関連性辞書情報は、前記メタ情報に含まれる単語と同義の他の単語を定義する情報を含むことを特徴とするデータ加工システム。
  6.  請求項2に記載のデータ加工システムであって、
     前記関連性辞書情報は、前記メタ情報に含まれる書き言葉の単語と同義の話し言葉の単語を定義する情報、および、前記メタ情報に含まれる話し言葉の単語と同義の書き言葉の単語を定義する情報を含むことを特徴とするデータ加工システム。
  7.  請求項2に記載のデータ加工システムであって、
     前記関連性辞書情報は、前記メタ情報に含まれる第1の言語の単語と同義の第2の言語の単語を定義する情報を含むことを特徴とするデータ加工システム。
  8.  請求項2に記載のデータ加工システムであって、
     前記メタ情報は、三項関係を含む情報であり、
     前記関連性辞書情報は、抽出されたメタ情報が三項関係に係る所定の条件に合致するときに、合致した三項関係を利用して新たな三項関係を生成するための規則を含むことを特徴とするデータ加工システム。
  9.  請求項2に記載のデータ加工システムであって、
     前記第1の種類のデータは、テキスト、表構造、音声、画像または文書のいずれかである第1のモーダルのデータであり、
     前記第2の種類のデータは、テキスト、表構造、音声、画像または文書のうち、前記第1のモーダルとは異なる第2のモーダルのデータであり、
     前記関連性辞書情報は、異なるモーダル間で置換される情報を定義する情報を含むことを特徴とするデータ加工システム。
  10.  一つ以上のプロセッサと、前記一つ以上のプロセッサに接続される一つ以上の記憶装置と、を有する計算機システムによるデータ加工方法であって、
     前記計算機システムは、複数の種類のデータからメタ情報を抽出する条件を定義するメタ情報抽出用辞書情報と、前記複数の種類のデータから抽出されたメタ情報を関連付ける条件を定義する関連性辞書情報と、を保持し、
     前記データ加工方法は、
     前記複数の種類のデータから、前記メタ情報抽出用辞書情報に基づいて前記メタ情報を抽出する第1手順と、
     入力されたデータからメタ情報を抽出する第2手順と、
     前記関連性辞書情報に基づいて、前記入力されたデータから抽出されたメタ情報と前記複数の種類のデータから抽出されたメタ情報とを関連付ける第3手順と、
     前記関連付けの結果に基づいて、前記複数の種類のデータと、前記入力されたデータと、それらのデータから抽出されたメタ情報と、のいずれかの組み合わせの関連を示す情報を出力する第4手順と、を含むことを特徴とするデータ加工方法。
  11.  請求項10に記載のデータ加工方法であって、
     前記複数の種類のデータは、第1の種類のデータと、第2の種類のデータと、を含み、
     前記メタ情報抽出用辞書情報は、前記第1の種類のデータから第1のメタ情報を抽出する条件を定義する第1のメタ情報抽出用辞書情報と、前記第2の種類のデータから第1のメタ情報を抽出する条件を定義する第2のメタ情報抽出用辞書情報と、を含み、
     前記関連性辞書情報は、メタ情報と置換可能な情報を定義する情報を含み、
     前記第1手順は、前記第1の種類のデータおよび前記第2の種類のデータから、それぞれ、前記第1のメタ情報抽出用辞書情報および前記第2のメタ情報抽出用辞書情報に基づいて、前記第1のメタ情報および前記第2のメタ情報を抽出する第5手順を含み、
     前記第2手順は、入力されたデータから第3のメタ情報を抽出する第6手順を含み、
     前記第3手順は、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報、前記第2のメタ情報と置換可能な情報、および前記第3のメタ情報と置換可能な情報の少なくとも一つを特定する第7手順と、
     前記第3のメタ情報または前記第3のメタ情報と置換可能な情報を用いて、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報、および、前記第2のメタ情報または前記第2のメタ情報と置換可能な情報を検索する第8手順と、
     前記検索の結果に基づいて、前記第3のメタ情報または前記第3のメタ情報と置換可能な情報と、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報と、前記第2のメタ情報または前記第2のメタ情報と置換可能な情報と、を関連付ける第9手順と、を含むことを特徴とするデータ加工方法。
  12.  請求項11に記載のデータ加工方法であって、
     前記第7手順は、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報を特定し、前記関連性辞書情報に基づいて、前記第2のメタ情報と置換可能な情報を特定し、前記特定された置換可能な情報と前記第1のメタ情報と前記第2のメタ情報とが所定の条件を満たすか否かの判定に基づいて、前記第1のメタ情報と前記第2のメタ情報とを関連付ける手順と、
     前記関連性辞書情報に基づいて、前記第3のメタ情報と置換可能な情報を特定する手順と、を含み、
     前記第9手順は、前記第8手順において前記第1のメタ情報または前記第1のメタ情報と置換可能な情報が検索結果として取得された場合、前記第3のメタ情報または前記第3のメタ情報と置換可能な情報と、前記第1のメタ情報または前記第1のメタ情報と置換可能な情報と、前記第1のメタ情報に関連付けられた前記第2のメタ情報または前記第2のメタ情報と置換可能な情報と、を関連付ける手順を含むことを特徴とするデータ加工方法。
  13.  請求項11に記載のデータ加工方法であって、
     前記第7手順は、
     前記関連性辞書情報に基づいて、前記第1のメタ情報と置換可能な情報を特定する手順、および、前記関連性辞書情報に基づいて、前記第2のメタ情報と置換可能な情報を特定する手順を含まず、
     前記関連性辞書情報に基づいて、前記第3のメタ情報と置換可能な情報を特定する手順を含むことを特徴とするデータ加工方法。
PCT/JP2012/084007 2012-12-28 2012-12-28 データ加工システムおよびデータ加工方法 Ceased WO2014102992A1 (ja)

Priority Applications (3)

Application Number Priority Date Filing Date Title
US14/649,762 US20150324436A1 (en) 2012-12-28 2012-12-28 Data processing system and data processing method
JP2014553983A JP5903171B2 (ja) 2012-12-28 2012-12-28 データ加工システムおよびデータ加工方法
PCT/JP2012/084007 WO2014102992A1 (ja) 2012-12-28 2012-12-28 データ加工システムおよびデータ加工方法

Applications Claiming Priority (1)

Application Number Priority Date Filing Date Title
PCT/JP2012/084007 WO2014102992A1 (ja) 2012-12-28 2012-12-28 データ加工システムおよびデータ加工方法

Publications (1)

Publication Number Publication Date
WO2014102992A1 true WO2014102992A1 (ja) 2014-07-03

Family

ID=51020143

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2012/084007 Ceased WO2014102992A1 (ja) 2012-12-28 2012-12-28 データ加工システムおよびデータ加工方法

Country Status (3)

Country Link
US (1) US20150324436A1 (ja)
JP (1) JP5903171B2 (ja)
WO (1) WO2014102992A1 (ja)

Families Citing this family (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11715151B2 (en) * 2020-01-31 2023-08-01 Walmart Apollo, Llc Systems and methods for retraining of machine learned systems
US20220351814A1 (en) * 2021-05-03 2022-11-03 Udo, LLC Stitching related healthcare data together

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008134954A (ja) * 2006-11-29 2008-06-12 Canon Inc 情報処理装置、その制御方法、及びプログラム
JP2008192102A (ja) * 2007-02-08 2008-08-21 Sony Computer Entertainment Inc メタデータ生成装置およびメタデータ生成方法
JP2008226110A (ja) * 2007-03-15 2008-09-25 Seiko Epson Corp 情報処理装置、情報処理方法および制御プログラム
JP2008236373A (ja) * 2007-03-20 2008-10-02 Nippon Hoso Kyokai <Nhk> メタ情報付加装置及びメタ情報付加プログラム

Family Cites Families (10)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US7117199B2 (en) * 2000-02-22 2006-10-03 Metacarta, Inc. Spatially coding and displaying information
JP2002221980A (ja) * 2001-01-25 2002-08-09 Oki Electric Ind Co Ltd テキスト音声変換装置
US20050209849A1 (en) * 2004-03-22 2005-09-22 Sony Corporation And Sony Electronics Inc. System and method for automatically cataloguing data by utilizing speech recognition procedures
US20060009966A1 (en) * 2004-07-12 2006-01-12 International Business Machines Corporation Method and system for extracting information from unstructured text using symbolic machine learning
US20070088549A1 (en) * 2005-10-14 2007-04-19 Microsoft Corporation Natural input of arbitrary text
US8615707B2 (en) * 2009-01-16 2013-12-24 Google Inc. Adding new attributes to a structured presentation
JP5525529B2 (ja) * 2009-08-04 2014-06-18 株式会社東芝 機械翻訳装置および翻訳プログラム
US8935259B2 (en) * 2011-06-20 2015-01-13 Google Inc Text suggestions for images
CN102955773B (zh) * 2011-08-31 2015-12-02 国际商业机器公司 用于在中文文档中识别化学名称的方法及系统
US9741344B2 (en) * 2014-10-20 2017-08-22 Vocalzoom Systems Ltd. System and method for operating devices using voice commands

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2008134954A (ja) * 2006-11-29 2008-06-12 Canon Inc 情報処理装置、その制御方法、及びプログラム
JP2008192102A (ja) * 2007-02-08 2008-08-21 Sony Computer Entertainment Inc メタデータ生成装置およびメタデータ生成方法
JP2008226110A (ja) * 2007-03-15 2008-09-25 Seiko Epson Corp 情報処理装置、情報処理方法および制御プログラム
JP2008236373A (ja) * 2007-03-20 2008-10-02 Nippon Hoso Kyokai <Nhk> メタ情報付加装置及びメタ情報付加プログラム

Also Published As

Publication number Publication date
JP5903171B2 (ja) 2016-04-13
JPWO2014102992A1 (ja) 2017-01-12
US20150324436A1 (en) 2015-11-12

Similar Documents

Publication Publication Date Title
US12032915B2 (en) Creating and interacting with data records having semantic vectors and natural language expressions produced by a machine-trained model
US9720944B2 (en) Method for facet searching and search suggestions
US10394851B2 (en) Methods and systems for mapping data items to sparse distributed representations
US12541543B2 (en) Large language model-based information retrieval for large datasets
US10073840B2 (en) Unsupervised relation detection model training
US10025819B2 (en) Generating a query statement based on unstructured input
CN108268600B (zh) 基于ai的非结构化数据管理方法及装置
CN108475289A (zh) 用于基于文档的内容向写作者建议内容的系统和方法
US20170177180A1 (en) Dynamic Highlighting of Text in Electronic Documents
US20160239504A1 (en) Method for entity enrichment of digital content to enable advanced search functionality in content management systems
JP6775935B2 (ja) 文書処理装置、方法、およびプログラム
CN106164890A (zh) 用于消除非结构化文本中的特征的歧义的方法
US20140195532A1 (en) Collecting digital assets to form a searchable repository
AU2016204573A1 (en) Common data repository for improving transactional efficiencies of user interactions with a computing device
US20120179709A1 (en) Apparatus, method and program product for searching document
US20230409624A1 (en) Multi-modal hierarchical semantic search engine
JPWO2016151690A1 (ja) 文書検索装置、方法及びプログラム
JP2018185716A (ja) データ処理システム、データ処理方法、およびデータ構造
US20210109960A1 (en) Electronic apparatus and controlling method thereof
NL2016846B1 (en) Computer implemented and computer controlled method, computer program product and platform for arranging data for processing and storage at a data storage engine.
JP7685921B2 (ja) 情報処理システム、情報処理方法、および情報処理プログラム
JP5903171B2 (ja) データ加工システムおよびデータ加工方法
US20160085850A1 (en) Knowledge brokering and knowledge campaigns
JP4900475B2 (ja) 電子文書管理装置及び電子文書管理プログラム
JP6905724B1 (ja) 情報提供システム及び情報提供方法

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application

Ref document number: 12890659

Country of ref document: EP

Kind code of ref document: A1

ENP Entry into the national phase

Ref document number: 2014553983

Country of ref document: JP

Kind code of ref document: A

WWE Wipo information: entry into national phase

Ref document number: 14649762

Country of ref document: US

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 12890659

Country of ref document: EP

Kind code of ref document: A1