WO2007105642A1 - 多義語による情報検索装置及びプログラム - Google Patents

多義語による情報検索装置及びプログラム Download PDF

Info

Publication number
WO2007105642A1
WO2007105642A1 PCT/JP2007/054692 JP2007054692W WO2007105642A1 WO 2007105642 A1 WO2007105642 A1 WO 2007105642A1 JP 2007054692 W JP2007054692 W JP 2007054692W WO 2007105642 A1 WO2007105642 A1 WO 2007105642A1
Authority
WO
WIPO (PCT)
Prior art keywords
extracted
articles
article
input
database
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2007/054692
Other languages
English (en)
French (fr)
Inventor
Masaki Murata
Kouichi Doi
Tomohiro Mitsumori
Yasushi Fukuda
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Nara Institute of Science and Technology NUC
National Institute of Information and Communications Technology
Original Assignee
Nara Institute of Science and Technology NUC
National Institute of Information and Communications Technology
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Nara Institute of Science and Technology NUC, National Institute of Information and Communications Technology filed Critical Nara Institute of Science and Technology NUC
Publication of WO2007105642A1 publication Critical patent/WO2007105642A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F16/00Information retrieval; Database structures therefor; File system structures therefor
    • G06F16/30Information retrieval; Database structures therefor; File system structures therefor of unstructured textual data
    • G06F16/35Clustering; Classification
    • G06F16/355Creation or modification of classes or clusters

Definitions

  • the present invention relates to an information retrieval apparatus and program using an ambiguous word that performs a search in consideration of the ambiguous word.
  • the word “WINS” has two terms: computer terms and horse racing terms. If you search only by entering “WINS”, search results related to computer terms and search results related to horse racing terms will be mixed and output. If the user wants search results only for articles related to computer terms, the above search results are inconvenient and need to be resolved.
  • Non-Patent Document 1 "Information search using location information and field information" Shingo Murata, Ma Aoi, Kiyotaka Uchimoto, Hiromi Osaku , Masao Uchiyama, Hitoshi Isahara, Natural Language Processing (Journal of the Language Processing Society) April 2000, No. 7, No. 2, P.141- P.160
  • An object of the present invention is to solve the above problems, perform a search in consideration of the ambiguity of words, and search (output) only necessary information.
  • FIG. 1 is an explanatory diagram of an information retrieval apparatus using a polysemy of the present invention.
  • 1 is an input section (input means)
  • 2 is a search extraction section (search extraction means)
  • 4 is a database (storage means)
  • 5 is an output section (output means).
  • the present invention has the following means in order to solve the above conventional problems.
  • the search extraction unit 2 extracts and outputs only the articles including the input keyword in the extracted similar articles. , Output in the order of the article power with the highest similarity to the article group B. Therefore, it is possible to reliably search for articles in the field entered using keywords with ambiguous terms.
  • Search / extraction means 2 for extracting expressions that appear biased in each cluster, and inquiry means for selecting expressions that appear unevenly in each cluster. The search / extraction means 2 is selected by the inquiry means.
  • the article of the cluster of the expressed expression is output. This makes it easy to search for articles in fields where you want to enter only keywords.
  • An input means 1 for inputting a keyword and a field, a database 4 for storing articles in each field, and an article including both the input keyword and field are extracted from the database 4.
  • the search extraction means 2 is a program for causing a computer to function. For this reason, by installing this program on a computer, it is possible to easily provide an information retrieval apparatus using a polysemy that can easily search for articles in a field in which only keywords are desired. Togashi.
  • the present invention has the following effects.
  • the search / extraction means extracts an article including the input keyword and field from the database, extracts a word group A that appears biased to the extracted article group, and includes the input keyword. Since the articles are output in the order of the article power that contains a large number of the word group A, it is possible to search for articles in the input field using keywords based on multiple terms.
  • FIG. 1 is an explanatory diagram of an information retrieval apparatus using a polysemy of the present invention.
  • FIG. 2 is a flowchart (1) of information retrieval using a polysemy of the present invention.
  • FIG. 3 is a flowchart (2) of information retrieval using a polysemy of the present invention.
  • FIG. 5 is a flowchart (3) of information retrieval using a polysemy of the present invention.
  • the information retrieval apparatus using ambiguous words performs retrieval in consideration of word ambiguity in information retrieval.
  • the word “WINS” has two terms: computer terminology and horse racing terminology. If you search only by entering "WINS”, search results related to computer terms and search results related to horse racing terms are mixed and output. If the user wants search results for only articles related to computer terms, the solution described below (Solutions 1 to 3) can be used.
  • FIG. 1 is an explanatory diagram of an information retrieval apparatus using ambiguous words.
  • an information retrieval device (system) using multiple terms includes an input unit (input unit) 1, a search extraction unit (search extraction unit) 2, a database (storage unit) 4, and an output unit (output unit) 5. It is provided.
  • the input unit 1 is an input means for inputting information such as keywords.
  • the search extraction unit 2 is a search extraction unit that performs word extraction, search processing, and the like.
  • Database 4 is a storage means for storing information (including information such as Web).
  • the output unit 5 is output means for outputting information by displaying and printing.
  • the user can enter the input form by specifying a field such as “keyword (field)”. For example, in the previous example, enter "WINS (computer)”.
  • FIG. 2 is a flowchart (1) of information retrieval using ambiguous words.
  • information retrieval using a multiple word Solution 1 will be described according to the processes S1 to S5 in FIG.
  • S1 The user inputs a keyword by designating a field using the input unit 1, and proceeds to processing S2.
  • S2 The search extraction unit 2 extracts an article including the keyword input from the database 4, and proceeds to processing S3.
  • S3 The search and extraction unit 2 extracts an article including the specified field from the extracted article group, and proceeds to processing S4.
  • S4 The search extraction unit 2 extracts a word group A that appears biased to the article group including the specified field from the article group including the input keyword, and proceeds to processing S5.
  • S5 The search extraction unit 2 outputs to the output unit 5 in the order of the article power including more word group A in the articles including the input keyword.
  • the article group can be used to extract word group A that appears biased to articles including computers.
  • C be a larger article group that contains article group B.
  • the article group C may be the whole database or a part thereof.
  • the article group includes 3 ⁇ 4 “WINS”.
  • Solution 1 described above may have other methods.
  • the database that does not extract word group A that appears biased in the articles that include the computer is not included.
  • the word group A that appears biased in the article group including the computer may be extracted from the entire article group, and processed using the extracted word group A. In that case, C is the entire database.
  • Appearance rate of A in C Number of occurrences of A in C Total number of words in ZC
  • Appearance rate of A in B Number of occurrences of A in B Total number of words in ZB
  • N be the number of occurrences of A in C.
  • N1 be the number of occurrences of A in B.
  • N2 N-N1.
  • N1 and N2 are not equivalent probabilities, that is, N1 is significantly larger than N2.
  • P1 is less than 5%, or 10% test, P1 is less than 10% is a criterion for determining whether it is significantly greater.
  • Words that appear to be biased in the article group B are those in which N1 is determined to be significantly larger than N2. In addition, the smaller P1, the more often the word appears in the article group B.
  • the number of occurrences of A in B is Nl, the total number of occurrences of words in B is Fl,
  • the number of occurrences of A that is in C but not in B is N2,
  • F2 be the total number of words that are in C but not in B.
  • R1 and R2 are more significant as the chi-square value is larger.
  • the chi-square value is greater than 3.84, it can be said that there is a significant difference of 5%, and the chi-square value is 6.63. If it is too large, it can be said that there is a significant difference of 1%.
  • test methods may be combined with the method of simply determining the appearance rate of A in B and the appearance rate of A in ZC.
  • W is a set of keywords entered by the user
  • N is the total number of documents
  • length is the length of article D
  • delta is the average length of articles
  • the length of the article uses the number of bytes of the article and the number of words included in the article.
  • E (t) 1 (keyword from the original search)
  • RatioC (t) is the appearance rate of t in article group B
  • RatioD (t) is the appearance rate of t in article group C
  • the score (D) is obtained by replacing the log (N / d w)) with the above equation, and the larger the value! /, the more the word group A is extracted.
  • the set W of words w to be added when score (D) is added is both the original keyword and the word group A. However, the original keyword and word group A should not overlap.
  • score (D) is added at the time of addition.
  • the set W of words w is only word group A. However, the original keyword and word group A should not overlap.
  • the user can enter the input form by specifying a field such as “keyword (field)”. For example, in the previous example, enter "WINS (computer)”.
  • a field such as “keyword (field)”. For example, in the previous example, enter "WINS (computer)”.
  • WINS computer
  • articles containing both “WINS” and the computer are first extracted. Then, similar articles in the article group B are extracted. In the similar articles, only articles that contain “WINS” are extracted and output as search results. At this time, articles with high similarity to article group B are output. This also seems to be able to extract articles in the computer-related field.
  • FIG. 3 is a flowchart (2) of information retrieval using multiple terms.
  • the process Sl l ⁇ in Fig. 3 In accordance with S14, explain information retrieval by using multiple meanings (Solution 2).
  • the search extraction unit 2 extracts articles including both the keyword and the field input from the database 4, and proceeds to processing S13.
  • S13 The search extraction unit 2 extracts similar articles in the extracted article group B, and proceeds to processing S14.
  • S14 The search extraction unit 2 extracts only the articles including the input keyword in the extracted similar articles, and outputs them as search results. At this time, it is output to the article power output unit 5 having a high similarity to the article group B.
  • _x, vector_y The value of _x, vector_y)) is obtained, and an article with a larger value may be determined as an article containing more word group A.
  • the word contained in the word group A is used as a vector (vector_x), and the word contained in the article is used as a vector (vector—y).
  • the similarity between the article group B and the article X includes the following methods.
  • the user inputs only “keyword”. For example, in the previous example, “WINS” is entered.
  • articles including “WINS” are extracted.
  • the articles are clustered. Extract expressions that appear biased in each cluster. For example, suppose that the expressions that are divided into two clusters and appear in each cluster are “computer” and “horse racing”, respectively. In that case, the user is inquired about whether it is related to “computer” or “horse racing”. Then, the user selects one of these. After the selection, the selected expression is processed as the input “field” in the same manner as in the above solutions 1 and 2, or the selected cluster is output as a search result.
  • FIG. 4 is an explanatory diagram of an information retrieval apparatus using a multiple word having an inquiry unit.
  • an information retrieval device (system) with a multiple meaning including an inquiry unit includes an input unit (input unit) 1, a search extraction unit (search extraction unit) 2, an inquiry unit (inquiry unit) 3, a database ( (Storage means) 4 and output unit (output means) 5 are provided.
  • the input unit 1 is an input means for inputting information such as keywords.
  • the search extraction unit 2 is a search extraction unit that performs word extraction, search processing, and the like.
  • the inquiry unit 3 is an inquiry means that asks the user for expressions (technical fields, etc.) that appear biased in the cluster, and makes selections by the user.
  • the database 4 is a storage means for storing information.
  • the output unit 5 is an output unit that outputs information by performing display and printing. [0081] (Description by flowchart)
  • FIG. 5 is a flowchart (3) of information retrieval using a polysemy.
  • information retrieval (solution 3) using a multiple meaning word having an inquiry part will be described according to the processes S21 to S26 in FIG.
  • the search extraction unit 2 extracts an article including the keyword input from the database 4, and proceeds to processing S23.
  • the search extraction unit 2 extracts expressions that appear unevenly in each cluster, and proceeds to processing S25.
  • S25 The inquiry unit 3 inquires the user so as to select an expression that appears biased in each cluster, and proceeds to processing S26.
  • the search extraction unit 2 outputs the articles of the selected cluster to the output unit 5.
  • Clusters and clusters are closest to each other.
  • the distance between cluster A and cluster B is the largest distance between the members of cluster A and cluster B, and the distance is the largest
  • the distance between cluster A and cluster B is the average of all cluster A member positions, and the average of all cluster B member positions is the single cluster position.
  • the average is the distance •
  • Ward method There is also a method called the Ward method. Hereinafter, the Ward method will be described.
  • x (i, j) is the position of the j-th member of the i-th cluster
  • ave— x (i) is the average of the positions of all members of the i-th cluster
  • the position of the member is the word taken from the article, the type of the word is taken as the dimension of the vector, and the value of the vector element of each word is set to the word frequency or the word 'idf (ie, tKw, D ) * log (N / dw) >> and the Okapi formula for that word (ie tl (w, D) / (ti (w, D) + length / delta) * log (N / dw)) Create and make it a member's position.
  • top-down clustering non-hierarchical clustering
  • clustering to a predetermined number k. Choose k members randomly, and use it as the center of the cluster. Each member becomes the closest cluster-centered member. The average of each member in the cluster is the center of each cluster. Each member becomes the closest cluster-centered member. In addition, the average of each member in the cluster The center of the raster. Repeat these. When the center of the cluster stops moving, it stops repeating. Or, repeat it for a predetermined number of times. The cluster is obtained by using the cluster center at the final cluster center. Each member is most recently a cluster-centered member.
  • clustering is performed. There are many other clustering methods that can be used.
  • the keyword given first may be plural, such as the force A B (B ′) C (C,) which is “WINS (computer)”. This means an AND search of word A, word B (but word B in the case of field B ') and word C (but word C in the case of field C').
  • Solution 3 is also possible. First, enter A, B, and C. Next, take out articles including A, B, and C. Clustering and outputting word Z that appears biased to each cluster. Simple The user can select a word and process the selected expression as the “field” of input in the same way as in solutions 1 and 2 above, or output the selected cluster as a search result.
  • Z1 co-occurs well with A
  • Z2 co-occurs with C
  • Z3 co-occurs with B
  • This display may take other forms as long as the relation between the input keyword and Zl, Z2,.
  • n (I ad— be I -n / 2) "2 / (a + b) / (c + d) / (a + c) / (b + d)
  • the process described as “taken out as the value is larger” can be taken out as “take out a value whose value is equal to or greater than the threshold value”.
  • the processing described as “take out a larger value in the order of the number greater than a predetermined value in order,” obtains a value obtained by multiplying the maximum value of the extracted value by a predetermined ratio, and “Take out the one with a value that is equal to or greater than the calculated value”.
  • these threshold values and predetermined values can be determined in advance, or the values can be appropriately changed and set by the user.
  • Input section (input means) 1, search extraction section (search extraction means) 2, question, matching section (question, matching means) 3, database (storage means) 4, output section (output means) 5, etc. are composed of programs It is executed by the main control unit (CPU) and is stored in the main memory.
  • This program is processed by a general computer (information processing apparatus).
  • This computer is composed of hardware such as an input device as input means such as a main control unit, main memory, file device, display device, and keyboard.
  • the program of the present invention is installed in this computer.
  • these programs are stored in a portable recording medium such as a hard disk or a magneto-optical disk, and the drive for accessing the recording medium provided in the computer is used.
  • It is installed in a file device provided in the computer via a device or a network such as a LAN. Then, the program steps necessary for the file device power processing are read out to the main memory and executed by the main control unit.

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Data Mining & Analysis (AREA)
  • Databases & Information Systems (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

 多義語によるキーワードを使用して入力した分野の記事を確実に検索する。  キーワードと分野を入力する入力手段1と、各分野の記事を格納するデータベース4と、前記入力したキーワードと分野を含む記事を前記データベース4から抽出し、該抽出した記事群に偏って出現する単語群Aを抽出し、前記入力したキーワードを含む記事の中で前記単語群Aを多く含む記事から順に出力する検索抽出手段3とを備える。

Description

明 細 書
多義語による情報検索装置及びプログラム
技術分野
[0001] 本発明は、言葉の多義性を考慮した検索を行う多義語による情報検索装置及びプ ログラムに関する。例えば、「WINS」という語は、コンピュータ用語と、競馬の用語の 二つがある。「WINS」とだけ入力して検索した場合は、コンピュータ用語に関連した 検索結果と、競馬の用語に関連する検索結果が混ざって出力される。もし、ユーザが コンピュータ用語に関連する記事だけの検索結果を欲しい場合は、上記の検索結果 では不便であるので、この問題を解決する必要がある。
背景技術
[0002] 従来、検索のためのキーワードを与えて情報検索を行う技術はあった (非特許文献 1参照)。しかし、検索の段階で、単語の多義を考慮した入力ができないものであった 非特許文献 1 : "位置情報と分野情報を用いた情報検索"村田真榭,馬青,内元清貴 ,小作浩美,内山将夫,井佐原均, 自然言語処理 (言語処理学会誌) 2000年 4月, 7卷, 2号, P.141〜 P.160
発明の開示
発明が解決しょうとする課題
[0003] 上記従来のキーワードを与えて情報検索を行う技術は、検索の段階で、単語の多 義を考慮した入力ができな力つたので、不必要な情報を検索して出力することがあつ た。
[0004] 本発明は上記問題点の解決を図り、言葉の多義性を考慮した検索を行い、必要な 情報のみを検索(出力)することを目的とする。
課題を解決するための手段
[0005] 図 1は本発明の多義語による情報検索装置の説明図である。図 1中、 1は入力部( 入力手段)、 2は検索抽出部 (検索抽出手段)、 4はデータベース (格納手段)、 5は出 力部(出力手段)である。 [0006] 本発明は、前記従来の課題を解決するため次のような手段を有する。
[0007] (1):キーワードと分野を入力する入力手段 1と、各分野の記事を格納するデータべ ース 4と、前記入力したキーワードと分野を含む記事を前記データベース 4から抽出 し、該抽出した記事群に偏って出現する単語群 Aを抽出し、前記入力したキーワード を含む記事の中で前記単語群 Aを多く含む記事力 順に出力する検索抽出手段 2と を備える。このため、多義語によるキーワードを使用して入力した分野の記事を検索 することができる。
[0008] (2):キーワードと分野を入力する入力手段 1と、各分野の記事を格納するデータべ ース 4と、前記入力したキーワードと分野を両方含む記事を前記データベース 4から 抽出し、該抽出した記事群 Bの類似記事を抽出し、該抽出した類似記事において、 前記入力したキーワードを含む記事のみを抽出して出力する検索抽出手段 2とを備 える。このため、多義語によるキーワードを使用して入力した分野の記事を検索する ことができる。
[0009] (3):前記(2)の多義語による情報検索装置において、前記検索抽出手段 2は、前 記抽出した類似記事において、前記入力したキーワードを含む記事のみを抽出して 出力する場合、前記記事群 Bとの類似度が高い記事力 順に出力する。このため、 多義語によるキーワードを使用して入力した分野の記事を確実に検索することができ る。
[0010] (4):キーワードを入力する入力手段 1と、各分野の記事を格納するデータベース 4 と、前記入力したキーワードを含む記事を前記データベース 4から抽出し、該抽出し た記事群をクラスタリングし、各クラスターで偏つて出現する表現を抽出する検索抽出 手段 2と、前記各クラスターで偏って出現する表現を選択する問い合わせ手段とを備 え、前記検索抽出手段 2は、前記問い合わせ手段で選択された表現のクラスターの 記事を出力する。このため、キーワードのみを入力してほしい分野の記事を容易に検 索することができる。
[0011] (5):前記(1)〜(3)の多義語による情報検索装置において、前記入力手段 1にキ 一ワードを入力し、前記検索抽出手段 2で前記入力したキーワードを含む記事を前 記データベース 4から抽出し、該抽出した記事群をクラスタリングし、各クラスターで偏 つて出現する表現を抽出し、前記各クラスターで偏って出現する表現を選択する問 V、合わせ手段を備え、前記問!、合わせ手段で選択された表現を前記入力手段 1に 入力される分野として用いる。このため、キーワードを入力して、ほしい分野の記事を 容易に検索することができる。
[0012] (6):キーワードと分野を入力する入力手段 1と、各分野の記事を格納するデータべ ース 4と、前記入力したキーワードと分野を含む記事を前記データベース 4から抽出 し、該抽出した記事群に偏って出現する単語群 Aを抽出し、前記入力したキーワード を含む記事の中で前記単語群 Aを多く含む記事力 順に出力する検索抽出手段 2と して、コンピュータを機能させるためのプログラムとする。このため、このプログラムをコ ンピュータにインストールすることで、多義語によるキーワードを使用して入力した分 野の記事を検索することができる多義語による情報検索装置を容易に提供すること ができる。
[0013] (7):キーワードと分野を入力する入力手段 1と、各分野の記事を格納するデータべ ース 4と、前記入力したキーワードと分野を両方含む記事を前記データベース 4から 抽出し、該抽出した記事群 Bの類似記事を抽出し、該抽出した類似記事において、 前記入力したキーワードを含む記事のみを抽出して出力する検索抽出手段 2として、 コンピュータを機能させるためのプログラムとする。このため、このプログラムをコンビ ユータにインストールすることで、多義語によるキーワードを使用して入力した分野の 記事を検索することができる多義語による情報検索装置を容易に提供することができ る。
[0014] (8):キーワードを入力する入力手段 1と、各分野の記事を格納するデータベース 4 と、前記入力したキーワードを含む記事を前記データベース 4から抽出し、該抽出し た記事群をクラスタリングし、各クラスターで偏つて出現する表現を抽出する検索抽出 手段 2と、前記各クラスターで偏って出現する表現を選択する問い合わせ手段と、前 記問い合わせ手段で選択された表現のクラスターの記事を出力する前記検索抽出 手段 2として、コンピュータを機能させるためのプログラムとする。このため、このプログ ラムをコンピュータにインストールすることで、キーワードのみを入力してほしい分野の 記事を容易に検索することができる多義語による情報検索装置を容易に提供するこ とがでさる。
発明の効果
[0015] 本発明によれば次のような効果がある。
[0016] (1):検索抽出手段で、入力したキーワードと分野を含む記事をデータベースから 抽出し、該抽出した記事群に偏って出現する単語群 Aを抽出し、前記入力したキー ワードを含む記事の中で前記単語群 Aを多く含む記事力 順に出力するため、多義 語によるキーワードを使用して入力した分野の記事を検索することができる。
[0017] (2):検索抽出手段で、入力したキーワードと分野を両方含む記事をデータベース 4から抽出し、該抽出した記事群 Bの類似記事を抽出し、該抽出した類似記事にお いて、前記入力したキーワードを含む記事のみを抽出して出力するため、多義語に よるキーワードを使用して入力した分野の記事を検索することができる。
[0018] (3):検索抽出手段で、抽出した類似記事において、入力したキーワードを含む記 事のみを抽出して出力する場合、記事群 Bとの類似度が高い記事力 順に出力する ため、多義語によるキーワードを使用して入力した分野の記事を確実に検索すること ができる。
[0019] (4):検索抽出手段で、入力したキーワードを含む記事をデータベース力 抽出し、 該抽出した記事群をクラスタリングし、各クラスターで偏つて出現する表現を抽出し、 問い合わせ手段で、前記各クラスターで偏って出現する表現を選択し、前記検索抽 出手段で、前記問 、合わせ手段で選択された表現のクラスターの記事を出力するた め、キーワードのみを入力してほしい分野の記事を容易に検索することができる。
[0020] (5):検索抽出手段で入力したキーワードを含む記事をデータベース力 抽出し、 該抽出した記事群をクラスタリングし、各クラスターで偏つて出現する表現を抽出し、 問!ヽ合わせ手段で前記各クラスターで偏って出現する表現を選択し、前記問 、合わ せ手段で選択された表現を前記入力手段に入力される分野として用いるため、キー ワードを入力して、ほし 、分野の記事を容易に検索することができる。
図面の簡単な説明
[0021] [図 1]本発明の多義語による情報検索装置の説明図である。
[図 2]本発明の多義語による情報検索のフローチャート(1)である。 [図 3]本発明の多義語による情報検索のフローチャート(2)である。
圆 4]本発明の問い合わせ部を備える多義語による情報検索装置の説明図である。
[図 5]本発明の多義語による情報検索のフローチャート(3)である。
符号の説明
[0022] 1 入力部 (入力手段)
2 検索抽出部 (検索抽出手段)
4 データベース (格納手段)
5 出力部(出力手段)
発明を実施するための最良の形態
[0023] 本発明の多義語による情報検索装置は、情報検索において言葉の多義性を考慮 した検索をするものである。例えば、「WINS」という語は、コンピュータ用語と、競馬 の用語の二つがある。「WINS」とだけ入力して検索した場合は、コンピュータ用語に 関連した検索結果と、競馬の用語に関連する検索結果が混ざって出力される。もし、 ユーザがコンピュータ用語に関連する記事だけの検索結果を欲しい場合は、以下で 説明する解決法 (解決方法 1〜3)で解決することができる。
[0024] (1):多義語による情報検索装置の説明
図 1は多義語による情報検索装置の説明図である。図 1において、多義語による情 報検索装置 (システム)には、入力部 (入力手段) 1、検索抽出部 (検索抽出手段) 2、 データベース (格納手段) 4、出力部(出力手段) 5が設けてある。
[0025] 入力部 1は、キーワード等の情報を入力する入力手段である。検索抽出部 2は、単 語の抽出、検索処理等を行う検索抽出手段である。データベース 4は、情報を格納 する格納手段 (Web等の情報も含む)である。出力部 5は、表示や印刷を行なって情 報を出力する出力手段である。
[0026] (2):多義語による情報検索の説明 1 (解決法 1)
ユーザが入力する形態を「キーワード (分野)」のように分野を指定して入力できるよ うにする。例えば、先の例だと、「WINS (コンピュータ)」と入力する。
[0027] この入力がなされると、まず、「WINS」を含む記事を抽出する。そして、その記事群 の中で、コンピュータを含む記事を抽出する。「WINS」を含む記事群の中で、コンビ ユータを含む記事群に偏って出現する単語群 Aを抽出する。 rwiNSjを含む記事の 中で単語群 Aをより多く含む記事力も順に出力する。単語群 Aはコンピュータ関連の 分野の記事に多く出現する表現で、そういう表現が多く出現する記事は、コンビユー タ関連の分野の記事と予想される。そういう記事を出力することで問題を解決する。
[0028] (フローチャートによる説明)
図 2は多義語による情報検索のフローチャート(1)である。以下、図 2の処理 S1〜S 5に従って、多義語による情報検索 (解決法 1)の説明をする。
[0029] S1 :入力部 1により、ユーザがキーワードを分野を指定して入力し、処理 S2に移る。
[0030] S2 :検索抽出部 2は、データベース 4から入力したキーワードを含む記事を抽出し、 処理 S3に移る。
[0031] S3 :検索抽出部 2は、抽出した記事群の中で、指定した分野を含む記事を抽出し、 処理 S4に移る。
[0032] S4 :検索抽出部 2は、入力したキーワードを含む記事群の中で、指定した分野を含 む記事群に偏って出現する単語群 Aを抽出し、処理 S5に移る。
[0033] S5 :検索抽出部 2は、入力したキーワードを含む記事の中で単語群 Aをより多く含 む記事力 順に出力部 5に出力する。
[0034] a)ある記事群 Bに偏って出現する単語群 Aの抽出方法の説明 1 (解決法 1)
例えば、コンピュータを含む記事群に偏って出現する単語群 Aを、抽出するときな どに使うことができる。記事群 Bを包含する、よりも大きい記事群を Cとする。ここで記 事群 Cはデータベース全体でもいいし、一部でもよい。上述の解決法 1にしたがえば 、 ¾「WINS」を含む記事群となる。
[0035] ただし、上述の解決法 1も他の方法がありえて、「WINS」を含む記事群の中で、コ ンピュータを含む記事群に偏って出現する単語群 Aを取り出すのではなぐデータべ ース全体の記事群の中で、コンピュータを含む記事群に偏って出現する単語群 Aを 取り出し、その取り出した単語群 Aを利用して処理してもよい。その場合は Cはデータ ベース全体となる。
[0036] 先ず、 C中の Aの出現率と B中の Aの出現率を求める。
[0037] C中の Aの出現率 =C中の Aの出現回数 ZC中の単語総数 B中の Aの出現率 =B中の Aの出現回数 ZB中の単語総数
次に、 B中の Aの出現率 ZC中の Aの出現率
を求めてこの値が大きいものほど、記事群 Bに偏って出現する単語とする。
[0038] b)ある記事群 Bに偏って出現する単語群 Aの抽出方法の説明 2
(有意差検定を利用する説明)
'二項検定の場合の説明
Aの Cでの出現数を Nとする。 Aの Bでの出現数を N1とする。
[0039] N2=N— N1とする。
[0040] Aが Cに現れたときにそれが B中に現れる確率を 0.5と仮定して、 Nの総出現のうち
、 N2回以下、 Aが Cに出現して Bに出現しな力つた確率を求める。
[0041] この確率は、
PI =∑ C(N1+N2,x) * 0.5 "(x) * 0.5 '(N1+N2— x)
(ただし、∑は、 x = 0力ら x = N2の禾口)
(ただし、 C(A,B)は、 A個の異なったものから B個のものを取り出す場合の数) (ただし、 'は、指数を意味する)
で表され、この確率の値が十分小さければ、 N1と N2は等価な確率でない、すなわち 、 N1が N2に比べて有意に大きいことと判断できる。
[0042] 5%検定なら
P1が 5%よりも小さいこと、 10%検定なら P1が 10%よりも小さいこと、が有意に大き いかどうかの判断基準になる。
[0043] N1が N2に比べて有意に大きいと判断されたものを記事群 Bに偏って出現する単 語とする。また、 P1が小さいものほど、記事群 Bによく偏って出現する単語とする。
[0044] 'カイ二乗検定の場合の説明
B中の Aの出現回数を Nl、 B中の単語の総出現数を Fl、
Cにあって Bにない、 Aの出現回数を N2、
Cにあって Bにない、単語の総出現数を F2とする。
[0045] N = N1 +N2として、
カイ二乗値 = (N * (Fl * (N2 - F2) - (N1 - Fl) * F2 )"2 )/((Fl + F2)*(N— (Fl + F 2)) * N1 * N2)
を求める。
[0046] そして、このカイ二乗値が大きいほど R1と R2は有意差があると言え、カイ二乗値が 3.84よりも大きいとき危険率 5%の有意差があると言え、カイ二乗値が 6.63よりも大 きいとき危険率 1%の有意差があると言える。
[0047] Nl > N2でかつ、カイ二乗値が大きいものほど、記事群 Bによく偏って出現する単 語とする。
[0048] ·比の検定、正確に言うと、比率の差の検定の説明
p = (F1+F2)/(N1+N2)
pi = Rl
p2 = R2
として、
Z = I pi - p2 I / sqrt ( p * (1 - p) * (1/Nl + 1/N2) )
を求め、(ただし sqrtはルートを意味する)そして、 Zが大きいほど、 R1と R2は有意 差があると言え、 Zが 1.96よりも大きいとき危険率 5%の有意差があると言え、 Zが 2. 58よりも大きいとき危険率 1%の有意差があると言える。
[0049] Nl > N2で、かつ、 Zが大きいものほど、記事群 Bによく偏って出現する単語とする。
[0050] これら三つの検定の方法と、先の単純に、 B中の Aの出現率 ZC中の Aの出現率を 求めて判定する方法を組み合わせてもよ 、。
[0051] 例えば、危険率 5%以上有意差があるもののうち、 B中の Aの出現率 ZC中の Aの 出現率、の値が大き!/ヽものほど記事群 Bによく偏って出現する単語とする。
[0052] c)単語群 Aをより多く含む記事の抽出方法の説明 (解決法 1)
情報検索の基礎知識として以下の式がある。ここで、 Score(D)が大きいものを取る。
[0053] (1)基本的な方法(TF · IDF法)の説明
score(D) =∑ ( tl(w,D) * log(N/dl(w)》
w £Wで加算
Wはユーザーが入力するキーワードの集合
t w,D)は文書 Dでの wの出現回数 d w)は全文書で Wが出現した文書の数
Nは文書の総数
score(D)が高い文書を検索結果として出力する。
[0054] (2) Robertsonらの Okapi weightingの説明
(文献)
村田真榭,馬青,内元清貴,小作浩美,内山将夫,井佐原均"位置情報と分野情 報を用いた情報検索"自然言語処理 (言語処理学会誌) 2000年 4月, 7卷, 2号, p. 141〜 P.160
の(1)式、が性能がよいことが知られている。これの式 (1)の∑で積を取る前のば項 と idf項の積が Okapiのウェイティング法になって、この値を単語の重みに使う。
[0055] Okapiの式なら
score(D) = ∑ ( tl(w,D)/(ti(w,D) + length/delta) * log(N/dl(w)) )
w £Wで加算
lengthは記事 Dの長さ、 deltaは記事の長さの平均、
記事の長さは、記事のバイト数、また、記事に含まれる単語数などを使う。
[0056] さらに、以下の情報検索を行うこともできる。
[0057] (Okapiの参考文献)
S. E. Robertson, b. Walker, b. Jones, M. M. Hancock— Beaulieu, and M. uatfor d Okapi at TREC— 3, TREC— 3, 1994
(SMARTの参考文献)
Amit Singhal AT&T at TREC— 6, TREC— 6, 1997
より高度な情報検索の方法として、 tf'idfを使うだけの式でなぐこれらの Okapiや S
MARTの式を用いてもよ!、。
[0058] これらの方法では、 tf'idfだけでなぐ記事の長さなども利用して、より高精度な情 報検索を行うことができる。
[0059] 今回の、単語群 Aをより多く含む記事の抽出方法では、さらに、 Rocchio's formula を使うことができる。
[0060] (文献) "]. J. Rocchio", "Relevance feedback in information retrieval", "The SMART retri eval System", "Edited by G. Salton", "Prentice Hall, Inc. , page 313-323〃, 1971 この方法は、 log(N/d w))のかわりに、
{E(t) + k_af * (RatioC(t) - RatioD(t))} *log(N/dl(w))
を使う。
[0061] E(t) = 1 (元の検索にあったキーワード)
= 0 (それ以外)
RatioC(t)は記事群 Bでの tの出現率
RatioD(t)は記事群 Cでの tの出現率
log(N/d w))を上式でおきかえた式で score(D)を求めて、その値が大き!/、ものほど、 単語群 Aをより多く含む記事として取り出すものである。
[0062] score(D)の∑の加算の際に足す単語 wの集合 Wは、元のキーワードと、単語群 Aの 両方とする。ただし、元のキーワードと、単語群 Aは重ならないようにする。
[0063] また、他の方法として、 score(D)の∑の加算の際に足す。単語 wの集合 Wは、単語 群 Aのみとする。ただし、元のキーワードと、単語群 Aは重ならないようにする。
[0064] ここでは roccioの式で複雑な方法をとつた力 単純に、単語群 Aの単語の出現回 数の和が大きいものほど、単語群 Aをより多く含む記事として取り出すようにしてもよ いし、また、単語群 Aの出現の異なりの大きいものほど、単語群 Aをより多く含む記事 として取り出すようにしてもょ 、。
[0065] (3):多義語による情報検索の説明 2 (解決法 2)
ユーザが入力する形態を「キーワード (分野)」のように分野を指定して入力できるよ うにする。例えば、先の例だと、「WINS (コンピュータ)」と入力する。この入力がなさ れると、まず、「WINS」とコンピュータの両方を含む記事を抽出する。そして、その記 事群 Bの類似記事を抽出する。その類似記事において、「WINS」を含む記事のみを 抽出し、それを検索結果として出力する。このとき記事群 Bとの類似度が高い記事か ら出力する。これも、コンピュータ関連の分野の記事を抽出できるものと思われる。
[0066] (フローチャートによる説明)
図 3は多義語による情報検索のフローチャート(2)である。以下、図 3の処理 Sl l〜 S14に従って、多義語による情報検索 (解決法 2)の説明をする。
[0067] S11 :入力部 1により、ユーザがキーワードを分野を指定して入力し、処理 S12に移 る。
[0068] S12 :検索抽出部 2は、データベース 4から入力したキーワードと分野を両方含む記 事を抽出し、処理 S13に移る。
[0069] S13 :検索抽出部 2は、抽出した記事群 Bの類似記事を抽出し、処理 S14に移る。
[0070] S14 :検索抽出部 2は、抽出した類似記事において、入力したキーワードを含む記 事のみを抽出し、それを検索結果として出力する。このとき記事群 Bとの類似度が高 い記事力 出力部 5に出力する。
[0071] a)記事群 Bの類似記事を抽出する方法の説明(解決法 2)
記事同士の類似度を定義する。この類似度は、 tf'idfや okapiや smartを使うとよい 。 tf'idfや okapiや smartなどにおける、記事 Dとクエリを比較する二つの記事 xと yと するとしてよい。そして、 x、 yの両方に含まれる単語^ wとするとよい。
[0072] 各単語を次元と、各単語のスコアを要素とするベクトルを作成し、記事 Xのベクトル を記事 Xに含まれる単語を使ってベクトル (vector— x)にし、また、記事 yのベクトルを 記事 yに含まれる単語を使ってベクトル (vector— y)にし、それらベクトルの余弦 (cos(v ector _x,vector_y))の値を記事の類似度としてもよい。各単語のスコアの算出には 、 tf'idfや okapiや smartを用いるとよい。それらの式の∑の後ろの部分の式がスコア の算出の式となる。その式の値が各単語のスコアとなる。
[0073] tf'idfだと t w,D) * log(N/d w))
okapiだと t w,D)/(t w,D) + length/delta) * log(N/dl(w))
がその式となる。
[0074] また、単語群 Aをより多く含む記事の抽出においてもこのベクトルの余弦 (cos(vector
_x,vector_y))の値を求め、この値が大きい記事ほど単語群 Aをより多く含む記事 と判断してもよい。この場合は、単語群 Aに含まれる単語を使ってベクトル (vector _x )にし、記事に含まれる単語を使ってベクトル (vector—y)にして求める。
[0075] 記事群 Bと記事 Xの類似度には、次の方法などがある。
[0076] ,記事群 Bのうち記事 Xと最も類似する記事と、記事 Xの類似度をその類似度とする 方法
•記事群 Bのうち記事 xと最も類似しない記事と、記事 xの類似度をその類似度とす る方法
•記事群 Bのすベての記事と記事 Xの類似度の平均をその類似度とする方法 他の方法でもよいが、このようにして、記事群 Bと記事 Xの類似度を求めて、その類 似度が大き 、ものを類似記事として取り出すことができる。
[0077] なお、他の方法としては、記事群 Bに偏って出現する単語を先の方法で取り出し、 そして、その単語も利用して、 Rocchio's formulaに基づく Score(D)を計算し、 Score( D)の大き 、ものを類似記事として取り出してもよ!、。
[0078] (4):多義語による情報検索の説明 3 (解決法 3)
ユーザは「キーワード」のみを入力する。例えば、先の例だと、「WINS」が入力され る。この入力がなされると、まず、「WINS」を含む記事を抽出する。そして、その記事 群をクラスタリングする。各クラスターで偏って出現する表現を抽出する。例えば、二 つのクラスターに分割され、それぞれのクラスターに偏って出現する表現が、それぞ れ、「コンピュータ」と「競馬」であったとする。その場合は、ユーザに、「コンピュータ」 と「競馬」のどちらに関連するかの問い合わせをする。そして、ユーザはこのいずれか を選択する。選択されたあとは、選択された表現を入力の「分野」として上記解決法 1 、 2と同様に処理するか、もしくは、選択されたクラスターを検索結果として出力する。
[0079] (問い合わせ部を備える多義語による情報検索装置の説明)
図 4は問い合わせ部を備える多義語による情報検索装置の説明図である。図 4に おいて、問い合わせ部を備える多義語による情報検索装置 (システム)には、入力部 (入力手段) 1、検索抽出部 (検索抽出手段) 2、問い合わせ部(問い合わせ手段) 3、 データベース (格納手段) 4、出力部(出力手段) 5が設けてある。
[0080] 入力部 1は、キーワード等の情報を入力する入力手段である。検索抽出部 2は、単 語の抽出、検索処理等を行う検索抽出手段である。問い合わせ部 3は、クラスターに 偏って出現する表現 (技術分野等)をユーザに問!、合わせ、ユーザが選択を行う問 い合わせ手段である。データベース 4は、情報を格納する格納手段である。出力部 5 は、表示や印刷を行なって情報を出力する出力手段である。 [0081] (フローチャートによる説明)
図 5は多義語による情報検索のフローチャート(3)である。以下、図 5の処理 S21〜 S26に従って、問い合わせ部を備える多義語による情報検索 (解決法 3)の説明をす る。
[0082] S21 :入力部 1により、ユーザがキーワードのみを入力し、処理 S22に移る。
[0083] S22 :検索抽出部 2は、データベース 4から入力したキーワードを含む記事を抽出し 、処理 S23に移る。
[0084] S23 :検索抽出部 2は、抽出した記事群をクラスタリングし、処理 S 24に移る。
[0085] S24 :検索抽出部 2は、各クラスターで偏って出現する表現を抽出し、処理 S25に 移る。
[0086] S25 :問い合わせ部 3は、各クラスターで偏って出現する表現の選択をするように、 ユーザに問い合わせ、処理 S26に移る。
[0087] S26 :検索抽出部 2は、選択されたクラスターの記事を出力部 5に出力する。
[0088] a)クラスタリングの説明(解決法 3)
クラスタリングにはさまざまな方法がある。一般的なものを以下に記述する。
[0089] (階層クラスタリング (ボトムアップクラスタリング)の説明)
最も近い成員同士をくつつけていき、クラスターを作る。クラスターとクラスター同士 も(クラスターと成員同士も)、最も近 、クラスター同士をくつつける。
クラスタ一間の距離の定義は様々あるので以下に説明する。
[0090] 'クラスター Aとクラスター Bの距離を、クラスター Aの成員とクラスター Bの成員の距 離の中で最も小さ ヽものをその距離とする方法
'クラスター Aとクラスター Bの距離を、クラスター Aの成員とクラスター Bの成員の距 離の中で最も大き 、ものをその距離とする方法
'クラスター Aとクラスター Bの距離を、すべてのクラスター Aの成員とクラスター Bの 成員の距離の平均をその距離とする方法
•クラスター Aとクラスター Bの距離を、すべてのクラスター Aの成員の位置の平均を そのクラスターの位置とし、すべてのクラスター Bの成員の位置の平均をそのクラスタ 一の位置とし、その位置同士の距離の平均をその距離とする方法 •ウォード法と呼ばれる方法もある。以下、ウォード法の説明をする。
[0091] W = ∑ ∑ (x(i,j) - ave _x(i)) " 2
Ίま指数を意味する。
[0092] 一つ目の∑は i=lから i=gまでの加算
二つ目の∑は j=lから j=niまでの加算
x(i,j)は i番目のクラスターの j番目の成員の位置
ave— x(i)は i番目のクラスターのすべての成員の位置の平均
クラスター同士をくつつけていくと、 Wの値が増加する力 ウォード法では、 Wの値が なるべく大きくならな 、ようにクラスター同士をくっつけて!/、く。
[0093] 成員の位置は、記事から単語を取り出し、その単語の種類をベクトルの次元とし、 各単語のベクトルの要素の値を、単語の頻度やその単語のば 'idf (すなわち、 tKw,D) * log(N/d w)》、その単語の Okapiの式(すなわち、 tl(w,D)/(ti(w,D) + length/delta) * log(N/d w)》としたベクトルを作成し、それをその成員の位置とする。
[0094] (トップダウンクラスタリング (非階層クラスタリング)の説明)
以下、トップダウンのクラスタリング (非階層クラスタリング)の方法を説明する。
[0095] (最大距離アルゴリズムの説明)
ある成員をとる。次にその成員と最も離れた成員をとる。これら成員をそれぞれのク ラスターの中心とする。それぞれのクラスター中心と、成員の距離の最小値を、各成 員の距離として、その距離が最も大きい成員をあらたなクラスターの中心とする。これ を繰り返す。あら力じめ定めた数のクラスターになったときに、繰り返しをやめる。また 、クラスタ一間の距離があら力じめ定めた数以下になると繰り返しをやめる。また、クラ スターの良さを AIC情報量基準などで評価してその値を利用して繰り返しをやめる方 法もある。各成員は、最も近いクラスター中心の成員となる。
[0096] (K平均法の説明)
あらカゝじめ定めた個数 k個にクラスタリングすることを考える。 k個成員をランダムに 選ぶ、それをクラスターの中心とする。各成員は最も近いクラスター中心の成員となる 。クラスター内の各成員の平均をそれぞれのクラスターの中心とする。各成員は最も 近いクラスター中心の成員となる。また、クラスター内の各成員の平均をそれぞれのク ラスターの中心とする。これらを繰り返す。そして、クラスターの中心が移動しなくなる と繰り返しをやめる。又は、あら力じめ定めた回数だけ繰り返してやめる。その最終的 なクラスター中心のときのクラスター中心を使ってクラスターを求める。各成員は最も 近 、クラスター中心の成員となる。
[0097] このようにして、クラスタリングをする。クラスタリングの方法は、これら以外にもたくさ んあるので、それらを利用してもよい。
[0098] b)各クラスターに偏って出現する表現の抽出の説明(解決法 3)
「ある記事群 Bに偏って出現する単語群 Aの抽出方法の説明 1 (解決法 1)」と同様 の方法で取り出すことが考えられ、そのようにしてもょ 、。
[0099] もっと単純には、各クラスターごとに、そのクラスターにしか出現しな力つた単語を頻 度順に並べて、各クラスターに偏って出現する表現として取り出しても良い。
[0100] (5) :複数のキーワードを用いる場合の説明
前記解決法 1、 2について、最初にあたえるキーワードは、「WINS (コンピュータ)」 になっている力 A B (B' ) C (C,)のように複数でもよい。これは、単語 Aと、単語 B (ただし、分野 B'の意味の場合の単語 B)と、単語 C (ただし、分野 C'の意味の場合 の単語 C)の AND検索を意味する。
[0101] a)解決法 1による説明
これを解決法 1で行う場合は、 A、 B、 Cを含む記事群 Xを取り出す。次に、記事群 X から B'、 C 'を含む記事群 X'を取り出す。記事群 Xのうち、記事群 X'に偏って出現す る単語群 Yを取り出す。そして、記事群 Xのうち、単語群 Yを多く含む記事を取り出し て出力する。
[0102] b)解決法 2による説明
これを解決法 2で行う場合は、 A、 B、 B'、 C、 C'を含む記事群 Xを取り出す。次に、 記事群 Xの類似記事を抽出する。類似記事において A、 B、 Cを含む記事を取り出し て出力する。
[0103] c)解決法 3による説明
解決法 3でもできる。まず、 A、 B、 Cを入力する。次に、 A、 B、 Cを含む記事群を取 り出す。クラスタリングして、各クラスターに偏って出現する単語 Zを出力する。その単 語をユーザーに選ばせて、選択された表現を入力の「分野」として上記解決法 1、 2と 同様に処理するか、もしくは、選択されたクラスターを検索結果として出力することが できる。
[0104] さらに、解決法 3では、各クラスターに偏って出現する単語群 Zを入力の A、 B、じと 対応づけて示すとよい。
[0105] 例えば、単語群 Zが頻度順に Zl, Z2, Z3,…としてあるとする。 Zl, Z2, Z3, ...を A、
B、 Cとよく共起するものと近づけて示してもよい。
[0106] Z1が Aとよく共起し、 Z2が Cとよく共起し、 Z3が Bとよく共起する場合
クラスター 1 A Zl、 B Z3、 C Z2
クラスター 2 のように表示して、 Zl, Z2, Z3, ..をユーザーに選ばせたり。クラスターをユーザに 選ばせる。なお、この表示は、入力キーワードと Zl, Z2,…の関連がわかるものなら ば他の形態でもよい。
[0107] Z1が Aとよく共起するかどうかは、次のものがある。
[0108] ·Ζ1と Aがともに出現する記事数が多いほど、よく共起するとするものとする。
[0109] ·前述の偏りの認識の方法を使い、 Z1を含む記事に、 Aがよく偏って出現すると判 断された場合、よく共起するとするものとする。
[0110] ·Ζ1と Aがともに出現する記事数を a、 Zlのみが出現する記事数を b、 Aのみが出現 する記事数を c、全記事数を dとして、
a
2a/(2a+b+c)
n(ad-bc) " 2/(a+b)/(c+d)/(a+c)/(b+d)
n( I ad— be I -n/2) " 2/(a+b)/(c+d)/(a+c)/(b+d)
log (an/(a+b)/(a+c))
(ad -bc)/((a+c)(b+d))"0.5
a log (an/(a+b)/(a+c)) + b log (bn/(a+b)/(b+d)) + c log (cn/(a+c)/(c+d)) + d log (dn /(b+d)/(c+d)) a/ (bc+ad)
a/ (ad- be)
a/b/c
などの値が大きいものを(これらのうちどれかの式を用いる)よく共起するとするもの とする。
[0111] など、 Z1が Aとよく共起するかどうかは、いろいろある。
[0112] なお、前記の実施の形態では、「値が大きいものほど取り出す」と記載した処理は「 値が閾値以上のものを取り出す」とすることができる。また、「値が大きいものを所定の 値の個数以上のものを大き 、順に取り出す」と記載した処理は「取り出されたものの 値の最大値に対して所定の割合をかけた値を求め、その求めた値以上の値を持つも のを取り出す」とすることができる。更に、これら閾値、所定の値を、あら力じめ定める ことも、適宜ユーザが値を変更、設定できることも可能である。
[0113] (9):プログラムインストールの説明
入力部 (入力手段) 1、検索抽出部 (検索抽出手段) 2、問 、合わせ部(問 、合わせ 手段) 3、データベース (格納手段) 4、出力部(出力手段) 5等は、プログラムで構成 でき、主制御部(CPU)が実行するものであり、主記憶に格納されているものである。 このプログラムは、一般的な、コンピュータ(情報処理装置)で処理されるものである。 このコンピュータは、主制御部、主記憶、ファイル装置、表示装置、キーボード等の入 力手段である入力装置などのハードウェアで構成されている。
[0114] このコンピュータに、本発明のプログラムをインストールする。このインストールは、フ 口ツビィ、光磁気ディスク等の可搬型の記録 (記憶)媒体に、これらのプログラムを記 憶させておき、コンピュータが備えている記録媒体に対して、アクセスするためのドラ イブ装置を介して、或いは、 LAN等のネットワークを介して、コンピュータに設けられ たファイル装置にインストールされる。そして、このファイル装置力 処理に必要なプ ログラムステップを主記憶に読み出し、主制御部が実行するものである。

Claims

請求の範囲
[1] キーワードと分野を入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードと分野を含む記事を前記データベースから抽出し、該抽出 した記事群に偏つて出現する単語群 Aを抽出し、前記入力したキーワードを含む記 事の中で前記単語群 Aを多く含む記事力 順に出力する検索抽出手段とを備えるこ とを特徴とした多義語による情報検索装置。
[2] キーワードと分野を入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードと分野を両方含む記事を前記データベースから抽出し、該 抽出した記事群 Bの類似記事を抽出し、該抽出した類似記事において、前記入力し たキーワードを含む記事のみを抽出して出力する検索抽出手段とを備えることを特 徴とした多義語による情報検索装置。
[3] 前記検索抽出手段は、前記抽出した類似記事において、前記入力したキーワードを 含む記事のみを抽出して出力する場合、前記記事群 Bとの類似度が高い記事力 順 に出力することを特徴とした請求項 2記載の多義語による情報検索装置。
[4] キーワードを入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードを含む記事を前記データベースから抽出し、該抽出した記 事群をクラスタリングし、各クラスターで偏って出現する表現を抽出する検索抽出手 段と、
前記各クラスターで偏って出現する表現を選択する問い合わせ手段とを備え、 前記検索抽出手段は、前記問い合わせ手段で選択された表現のクラスターの記事 を出力することを特徴とした多義語による情報検索装置。
[5] 前記入力手段にキーワードを入力し、前記検索抽出手段で前記入力したキーワード を含む記事を前記データベースから抽出し、該抽出した記事群をクラスタリングし、各 クラスターで偏って出現する表現を抽出し、
前記各クラスターで偏って出現する表現を選択する問い合わせ手段を備え、 前記問い合わせ手段で選択された表現を前記入力手段に入力される分野として用 いることを特徴とした請求項 1〜3のいずれかに記載の多義語による情報検索装置。
[6] キーワードと分野を入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードと分野を含む記事を前記データベースから抽出し、該抽出 した記事群に偏つて出現する単語群 Aを抽出し、前記入力したキーワードを含む記 事の中で前記単語群 Aを多く含む記事力 順に出力する検索抽出手段として、 コンピュータを機能させるためのプログラム。
[7] キーワードと分野を入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードと分野を両方含む記事を前記データベースから抽出し、該 抽出した記事群 Bの類似記事を抽出し、該抽出した類似記事において、前記入力し たキーワードを含む記事のみを抽出して出力する検索抽出手段として、 コンピュータを機能させるためのプログラム。
[8] キーワードを入力する入力手段と、
各分野の記事を格納するデータベースと、
前記入力したキーワードを含む記事を前記データベースから抽出し、該抽出した記 事群をクラスタリングし、各クラスターで偏って出現する表現を抽出する検索抽出手 段と、
前記各クラスターで偏って出現する表現を選択する問い合わせ手段と、 前記問い合わせ手段で選択された表現のクラスターの記事を出力する前記検索抽 出手段として、
コンピュータを機能させるためのプログラム。
PCT/JP2007/054692 2006-03-10 2007-03-09 多義語による情報検索装置及びプログラム Ceased WO2007105642A1 (ja)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2006065291A JP4857448B2 (ja) 2006-03-10 2006-03-10 多義語による情報検索装置及びプログラム
JP2006-065291 2006-03-10

Publications (1)

Publication Number Publication Date
WO2007105642A1 true WO2007105642A1 (ja) 2007-09-20

Family

ID=38509465

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2007/054692 Ceased WO2007105642A1 (ja) 2006-03-10 2007-03-09 多義語による情報検索装置及びプログラム

Country Status (3)

Country Link
JP (1) JP4857448B2 (ja)
CN (1) CN101405725A (ja)
WO (1) WO2007105642A1 (ja)

Families Citing this family (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP5388038B2 (ja) * 2009-12-28 2014-01-15 独立行政法人情報通信研究機構 文書要約装置、文書処理装置、及びプログラム
WO2011153708A1 (zh) * 2010-06-11 2011-12-15 上海坦瑞信息技术有限公司 一种基于领域概念的信息搜索方法
EP2635981A4 (en) * 2010-11-01 2016-10-26 Microsoft Technology Licensing Llc IMAGE SEARCH
CN102033961A (zh) * 2010-12-31 2011-04-27 百度在线网络技术(北京)有限公司 一种开放式知识共享平台及其多义词展现方法
JP5972096B2 (ja) * 2012-08-08 2016-08-17 Kddi株式会社 コンテンツに関する投稿を抽出する装置、方法およびプログラム
JP6007088B2 (ja) * 2012-12-05 2016-10-12 Kddi株式会社 大量のコメント文章を用いた質問回答プログラム、サーバ及び方法
CN104008098B (zh) * 2013-02-21 2018-09-18 腾讯科技(深圳)有限公司 基于多义性关键词的文本过滤方法及装置
CN108920467B (zh) * 2018-08-01 2021-04-27 北京三快在线科技有限公司 多义词词义学习方法及装置、搜索结果显示方法

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH1145274A (ja) * 1997-07-28 1999-02-16 Just Syst Corp 単語間の共起性を用いたキーワードの拡張方法およびその方法の各工程をコンピュータに実行させるためのプログラムを記録したコンピュータ読み取り可能な記録媒体
JP2000250925A (ja) * 1999-02-26 2000-09-14 Matsushita Electric Ind Co Ltd 文書検索・分類方法および装置
JP2003208447A (ja) * 2002-01-11 2003-07-25 Nippon Telegr & Teleph Corp <Ntt> 文書検索装置、文書検索方法、文書検索プログラム及び文書検索プログラムを記録した媒体
JP2004086635A (ja) * 2002-08-27 2004-03-18 Nri & Ncc Co Ltd 概念検索システム、概念検索方法およびコンピュータプログラム

Family Cites Families (11)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2542464B2 (ja) * 1991-09-20 1996-10-09 日本電信電話株式会社 文書検索装置
JPH0676004A (ja) * 1992-07-06 1994-03-18 Nec Corp データベース検索解表示装置
JP4075094B2 (ja) * 1997-04-09 2008-04-16 松下電器産業株式会社 情報分類装置
JP2000148764A (ja) * 1998-11-05 2000-05-30 Fujitsu Ltd クラスタリングを用いた検索質問展開処理装置,検索質問展開処理方法および検索質問展開処理用プログラム記録媒体
JP2001005830A (ja) * 1999-06-23 2001-01-12 Canon Inc 情報処理装置及びその方法、コンピュータ可読メモリ
JP2002132824A (ja) * 2000-10-26 2002-05-10 Seiko Epson Corp 情報検索方法および情報検索システム
JP3862059B2 (ja) * 2001-01-22 2006-12-27 Kddi株式会社 検索式拡張方法および検索システム
JP4092933B2 (ja) * 2002-03-20 2008-05-28 富士ゼロックス株式会社 文書情報検索装置及び文書情報検索プログラム
JP2004295797A (ja) * 2003-03-28 2004-10-21 Oki Electric Ind Co Ltd 情報検索装置
JP4344207B2 (ja) * 2003-09-19 2009-10-14 株式会社リコー 文書検索装置、文書検索方法、文書検索プログラム、および記録媒体
JP4569179B2 (ja) * 2004-06-03 2010-10-27 富士ゼロックス株式会社 ドキュメント検索装置

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH1145274A (ja) * 1997-07-28 1999-02-16 Just Syst Corp 単語間の共起性を用いたキーワードの拡張方法およびその方法の各工程をコンピュータに実行させるためのプログラムを記録したコンピュータ読み取り可能な記録媒体
JP2000250925A (ja) * 1999-02-26 2000-09-14 Matsushita Electric Ind Co Ltd 文書検索・分類方法および装置
JP2003208447A (ja) * 2002-01-11 2003-07-25 Nippon Telegr & Teleph Corp <Ntt> 文書検索装置、文書検索方法、文書検索プログラム及び文書検索プログラムを記録した媒体
JP2004086635A (ja) * 2002-08-27 2004-03-18 Nri & Ncc Co Ltd 概念検索システム、概念検索方法およびコンピュータプログラム

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
MURATA M. ET AL.: "Ichi Joho to Bun'ya Joho o Mochiita Joho Kensaku", JOURNAL OF NATURAL LANGUAGE PROCESSING, vol. 7, no. 2, 10 April 2000 (2000-04-10), pages 141 - 160, XP003017079 *
YAMAZAKI T.: "Ayamari Kudogata Gakushu to Thesaurus o Mochiita Bunsho Jido Bunrui", IEICE TECHNICAL REPORT, vol. 97, no. 2000, 27 July 1997 (1997-07-27), pages 19 - 26, XP003017080 *

Also Published As

Publication number Publication date
JP4857448B2 (ja) 2012-01-18
CN101405725A (zh) 2009-04-08
JP2007241794A (ja) 2007-09-20

Similar Documents

Publication Publication Date Title
Zhong et al. Effective pattern discovery for text mining
Nallapati Discriminative models for information retrieval
Chirita et al. P-tag: large scale automatic generation of personalized annotation tags for the web
Uğuz A two-stage feature selection method for text categorization by using information gain, principal component analysis and genetic algorithm
US8117185B2 (en) Media discovery and playlist generation
Song et al. Overview of the NTCIR-9 INTENT Task.
Alguliev et al. Evolutionary algorithm for extractive text summarization
Trappey et al. An R&D knowledge management method for patent document summarization
US8346800B2 (en) Content-based information retrieval
CN110023924A (zh) 用于语义搜索的设备和方法
WO2007105642A1 (ja) 多義語による情報検索装置及びプログラム
Wang et al. Targeted disambiguation of ad-hoc, homogeneous sets of named entities
JP2005302042A (ja) マルチセンスクエリについての関連語提案
Clinchant et al. Xrce’s participation in wikipedia retrieval, medical image modality classification and ad-hoc retrieval tasks of imageclef 2010
Song et al. An effective query recommendation approach using semantic strategies for intelligent information retrieval
CN118467708B (zh) 基于协同增强的词项级查询扩展方法
US9164981B2 (en) Information processing apparatus, information processing method, and program
Bian et al. Ranking specialization for web search: a divide-and-conquer approach by using topical ranksvm
Sutanto et al. The ranking based constrained document clustering method and its application to social event detection
AL-Khassawneh et al. Improving triangle-graph based text summarization using hybrid similarity function
Mei et al. Semantic annotation of frequent patterns
Al-Shboul et al. Query phrase expansion using wikipedia in patent class search
Zhang et al. A preprocessing framework and approach for web applications
Li et al. Keyphrase Extraction and Grouping Based on Association Rules.
Verberne et al. Author-topic profiles for academic search

Legal Events

Date Code Title Description
121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 200780008681.4

Country of ref document: CN

NENP Non-entry into the national phase

Ref country code: DE

122 Ep: pct application non-entry in european phase

Ref document number: 07738177

Country of ref document: EP

Kind code of ref document: A1