WO2007010836A1 - コミュニティ特有表現検出装置及び方法 - Google Patents

コミュニティ特有表現検出装置及び方法 Download PDF

Info

Publication number
WO2007010836A1
WO2007010836A1 PCT/JP2006/314000 JP2006314000W WO2007010836A1 WO 2007010836 A1 WO2007010836 A1 WO 2007010836A1 JP 2006314000 W JP2006314000 W JP 2006314000W WO 2007010836 A1 WO2007010836 A1 WO 2007010836A1
Authority
WO
WIPO (PCT)
Prior art keywords
word
community
gram
significance
selecting
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Ceased
Application number
PCT/JP2006/314000
Other languages
English (en)
French (fr)
Inventor
Hiromi Oda
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Hewlett Packard Development Co LP
Original Assignee
Hewlett Packard Development Co LP
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Hewlett Packard Development Co LP filed Critical Hewlett Packard Development Co LP
Priority to CN2006800258021A priority Critical patent/CN101223521B/zh
Priority to JP2007525983A priority patent/JPWO2007010836A1/ja
Priority to US11/990,495 priority patent/US20100076745A1/en
Priority to DE112006001822T priority patent/DE112006001822T5/de
Publication of WO2007010836A1 publication Critical patent/WO2007010836A1/ja
Anticipated expiration legal-status Critical
Ceased legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/237Lexical tools
    • G06F40/242Dictionaries
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • G06F40/279Recognition of textual entities
    • G06F40/289Phrasal analysis, e.g. finite state techniques or chunking

Definitions

  • the present invention relates to an apparatus and method for detecting a community-specific expression from expressions used in a community based on word formation theory.
  • Patent Document 1 JP 2002-297589 “Unknown word collection method”
  • Patent Document 2 JP-A-5-113997 “Dictionary Data Collection Device”
  • Patent Document 3 JP 2004-265440 “Unknown Word Registration Device and Method and Storage Medium”
  • Patent Document 4 JP 2005-309853 “Vocabulary Conversion Method Between Professional Description and Non-Professional Description 'Program' System”
  • Non-patent document 1 Yuji Nakagawa, Yasuaki Yumoto, & Nada Nada (2003). Extraction of specialized terms based on appearance frequency and connection frequency. Natural language processing, 10 (1), 27-45.
  • Non-Patent Literature 2 Zhaoqing University, & Fuyue Fumane (2004). Basic Research for Identifying New Words Important in Specialized Fields. Proc. Of the 10th Annual Conference of the Language Processing Society, (pp. 189 -191).
  • Non-Patent Document 3 Satoshi Fujii, Katsunobu Ito, Tomoaki Akiba (2003), IPA Unexplored Software Creation Project “CYCLONE: Building the Strongest Dictionary Site”, www.ipa.go.jp/about/news/event/ pdf / 29A7_f ujii.pdf
  • Non-patent document 4 Akihiko Yonekawa (1998) “Science of youth language” Tokyo: Meiji Shoin
  • Patent Document 1 Japanese Patent Application Laid-Open No. 2002-297589 “Unknown Word Collection Method”
  • Patent Document 3 Japanese Patent Application Laid-Open No. 2004-265440 “Unknown Word Registration Device”
  • Patent Document 1 Japanese Patent Application Laid-Open No. 2002-297589 “Unknown Word Collection Method”
  • Patent Document 1 This method also has the same power. Basically, many things that are not registered in the dictionary are collected by human stakes. In the detection of these unknown words, the target is almost limited to nouns, and rarely focus on the problem of collecting truly new expressions.
  • Non-patent Document 4 In sociolinguistics, there is a field that collects and analyzes “young people” used by high school students and university students (Non-patent Document 4). Existing research on community-specific expressions seems to be close to the present invention, but in the field of sociolinguistics, it has been proposed that a method should be proposed for regularly collecting youth and buzzwords.
  • Document gathering power used in a given community with the following means (a) to (d) A device that searches for expressions unique to a given community,
  • the apparatus according to (1) further comprising means for collecting the document set by performing a data search using a term included in a predetermined term list as a keyword.
  • the means for extracting the n-gram collocation uses a document used in a plurality of communities, and calculates the significance of the n-gram collocation used in the predetermined community and the n-gram collocation used in other communities.
  • a method for retrieving an expression specific to a given community from a set of documents used in the given community comprising the following steps (a) to (d):
  • the program according to (6) further comprising means for collecting the document set by searching data using a term included in a predetermined term list as a keyword.
  • the invention of the present application is an extension of the language between main parts of speech and can be applied to other languages.
  • the expression “He 747'ed to Chicago.” Is possible. This is a verbal version of the aircraft model. Also, "The web-logging is becoming a social phenomenon.” This is an example of a noun verb.
  • FIG. 1 shows an example of a system when the present invention is implemented.
  • a user PC 110 Connected to the network 140 are a user PC 110, a site server (1) 120, a site server (2) 130, and the like.
  • the site server (1) 120, site server (2) 130, etc. connected to the network 140 are accessed, and necessary information is acquired using a search tool or the like.
  • the present invention shows a search on the Internet as an embodiment, the present invention is not limited to this, and any other method can be applied as long as the system can search information.
  • the acquired information can be processed by a computer program on the user PC to obtain the desired result.
  • FIG. 2 shows a user PC that implements part of the present invention.
  • the housing 200 includes a storage device 210, a main memory 220, an output device 230, a central control device (CPU) 240, an operation device 250, and a network 1/0260.
  • the user operates the operation device 250 and obtains necessary information from each site on the Internet through the network I / O.
  • the central controller 240 downloads the document processing program stored in the storage device 210 to the memory, performs predetermined data processing using information retrieved from the Internet, and displays the result on the output device 230. .
  • FIG. 3 shows a block diagram of a community specific expression detection apparatus according to the present invention.
  • 3 10 is a community document search unit
  • 314 is a website
  • 316 is a term list storage unit
  • 320 is a document processing unit
  • 330 is an n-gram collocation extraction unit
  • 335 is a significance determination unit
  • 340 is a word base selection unit
  • 350 is The left and right extension part of the word base
  • 354 is the left extension rule storage part
  • 356 is the right extension rule storage part
  • 360 is the new expression selection part
  • 365 is the language rule storage part
  • 370 is the output part. Details of these will be described below.
  • Step 410 Collect documents for community use
  • Step 420 n-gram collocation extraction
  • Step 430 Selecting the core element (word base) of the new expression
  • Step 440 Select extended word base
  • Step 450 New expression selection
  • Step 510 Get candidate documents by specifying terms
  • Step 520 Preprocessing candidate documents
  • Step 530 Remove noise document
  • Step 540 Need to search for other community documents
  • Step 510 Acquisition of candidate documents
  • a term list including a predetermined term is used to collect documents used by parties in a predetermined community.
  • the term list is stored in the term list storage unit (Fig. 3: 316).
  • the term list is a set of terms that become keywords in one community. For example, if “wine lovers” is selected as one community, the component of the term list is “wine brands”. According to the brands listed in the wine terminology, use the Internet search tool to collect information about the wine ( Figure 3: 314). Here, brands such as “Hauslese”, “Chateau Kyule Bonn”, “Chateau Margoichi”, “Vine Santo Toscano” and the like can be designated. Candidate documents are searched from the database using this term as a keyword. Any database can be used as long as such information is stored in the database, but in this embodiment, a method for searching candidate documents using an Internet search engine will be described. [0017] (1 2) Step 520: Preprocessing of candidate document
  • the web page information-powered document is first extracted and analyzed.
  • segmentation is performed to extract content words, particles, auxiliary verbs, etc., and feature values representing the characteristics of these documents are obtained.
  • feature values representing the characteristics of these documents are obtained.
  • noise documents are removed as follows.
  • Documents that automatically collect this information from Internet web pages contain a variety of information and are often not available as they are.
  • documents corresponding to garbage documents, list documents, and diary documents are removed from these documents as noise documents.
  • a document that satisfies all the conditions such as a document with a small number of content words or a document with a low proper noun ratio.
  • the number of content words is the number of content words contained in a document described on one web page.
  • Content words are words that correspond to nouns, verbs, adjectives, and adverbs, excluding particles and auxiliary verbs.
  • the proper nouns mentioned here are nouns that are generally recognized as proper nouns.
  • the proper noun ratio is the ratio between the number of proper nouns appearing on one web page and the number of content words.
  • a document that satisfies all of the conditions such as a document having a high proper noun ratio, a document having a low correlation coefficient between the content word and the particle 'auxiliary verb', etc. is defined as a list information document. This is a document where information about objects in a certain area is stored as a simple list on an Internet site.
  • a document to be defined is defined as a diary document. These are so-called These are documents that mainly contain other information such as documents used as personal diary writing sites and sites related to department stores. Based on the above definition, garbage documents, list documents, and diary documents are removed as noise documents.
  • Step 540 Necessity of Search for Other Community Documents
  • step 510 From step 510 to step 530, a set of documents used in a predetermined community is collected.
  • step 540 a collection of documents used by other communities is collected as well.
  • n-gram collocations word-level n-gram collocations (n-gram collocations) using statistical methods and those that appear significantly when used in a specific community. These are called community-specific collocations. These details will be described.
  • An n-gram collocation is a sequence of one or more words: a unigram for one word, a bigram for two words, a trigram for three words. It is called (Tri-gram).
  • Tri-gram bigrams and trigrams are used (FIG. 3: 330).
  • the sample ratio is a ratio obtained from actual data
  • pi and p2 are sample ratios.
  • n-gram collocation W means to test whether it appears significantly biased towards the document in dl. Yes (one-sided test).
  • null hypothesis the null hypothesis and the alternative hypothesis are as follows.
  • a list of 2 grams and 3 grams appearing characteristically in a document set used by wine lovers and a document set used by sake lovers is extracted, and a Z test is performed.
  • n-grams with a Z value of 1.65 or more are selected from a set of documents used by wine lovers.
  • the n-gram extracted by the above method Take out the core element ( Figure 3: 340). To do this, break the n-gram chain for the time being and make a list of all the elements (morphemes) that occur there. From there, exclude those that are not likely to be core. Here, there is a function such as a particle, an auxiliary verb, a conjunction, a conjugation ending, and a break element such as “,”, “.”, “?”, Etc. as those that are not likely to be the core. Also excluded are “one hiragana character” and “one katakana character”. This creates a list of elements (the core list) that can be the core of the new expression.
  • each word base candidate it is determined whether it is necessary to incorporate the preceding and succeeding elements based on the collocation pattern distribution (Fig. 3: 350).
  • Z [X] is the Z value of the n-gram word group that we are currently focusing on.
  • X be the core element
  • [X + 1] be the element expanded by one word
  • [X + 2] be the element expanded by two words.
  • AvgZ ([X] [X + l]) is the word of all (n + 1) grams corresponding to [X] [X + 1] when expanded from the n-gram word base to the right It is the average of the base Z values (0 ⁇ Z;).
  • Equation 6 is defined by taking the logarithm of Z.
  • (ii) LZ> first threshold If it satisfies, it is selected as a candidate to expand to [X + 1] (610, 620, 650).
  • the first threshold value is 5.0 in this embodiment, and Z ([X], [X + 1]) is represented by ([X], [X + 1]) (n + 1) Gram word base Z value of AvgZ ([X], [X + l], [X + 2]) is all (n + 2) grams corresponding to [X], [X + 1], [X + 2] This is the average of the Z values.
  • the first threshold for LZ used in the first condition is set high. If this value is high, it will be judged that it can be recognized as a new expression enough even by judgment based on the value of Z. Therefore, it is selected as a possibility of new expression regardless of the value of Jratio (described later). To do.
  • condition (ie) both conditions (i) and (ii) are met, it is selected as an expanded word candidate (650). If condition (i) is not met! /, It is not selected as a candidate for expansion (660). If the condition (i) is satisfied but the condition (ii) is not satisfied, the determination is made based on the second condition shown below (630, 640).
  • the second threshold value for LZ used in the second condition is set to 3.0 in the example, and only when LZ is larger than this value and Jratio is 0.1 or more, new expression is possible. It is determined that there is sex.
  • Nail is the number of (n + 2) grams corresponding to the target [X + 2].
  • the elements of [X + 2], that is, “ga” and “ha” are called kOne elements. If there are multiple kOne elements as in this example, the average value of these Z values is calculated. In this case, since both are 2.00, the average value is 2.00.
  • this kOne element is a “break element” indicating a break.
  • a break element indicating a break.
  • a grammatical break is shown.
  • Jratio The proportion of kOne elements that are break elements is called Jratio.
  • the left extension rule is explained using an example. Explain that [receiving] (Z value is 73.01) selected as a word base is extended to the left.
  • Nounization and Examples include “base + suffix”, “verb conjunctive nounization”, “compound noun”, and the like. In each case, it is necessary to confirm the key to satisfy the rules for Japanese.
  • the present invention can be applied not only to Japanese but also to foreign languages. I will explain using English as an example. Something that is used in English as a part of speech other than the original noun may be used as a noun. For example, it is made a noun by adding the following suffix. “Ness”: pleasantness, ugliness
  • Verbification rules (step 720) Those that match the verbalization rules are also selected as candidates for word base expansion. Examples of verbs include “noun + do” and “general use of verb”. It is necessary to confirm whether the candidate selected for expansion satisfies the Japanese rules.
  • a noun is combined with a verbal suffix such as “S”, “Buru”, or its conjugation, it is selected as a candidate for verbal expansion of the word base. For example, if “tea” is added to “tea” and “tea is made”, “beauty” is added to “beauty” by adding “bu”.
  • An expanded word base is also selected as a candidate for expansion of the word base even if it is a general verb usage form excluding the form of “noun + verbal suffix”.
  • verbs are added to the nouns and converted into verbs: “Demo, not demo, if demo”.
  • new L ⁇ verbs such as “Gevaru, Hamoru, Tsumoru, Darguru” can be created in this way.
  • the present invention can be applied not only to Japanese but also to foreign languages. I will explain using English as an example. Something that is originally used as a noun in English may be used as a verb. Are you googling?
  • Step 710 to Step 740 If any of the above conditions from Step 710 to Step 740 is satisfied, it is selected as a candidate for expansion of the word base (760). If neither condition is met, it is not selected as a candidate for expansion of the word base (750).
  • the LZ value is 3.01.
  • expanded compound nouns include:
  • FIG. 1 is a diagram showing an example of a system for carrying out the present invention.
  • FIG. 2 is a block diagram of a PC that implements part of the present invention.
  • FIG. 3 is a block diagram of a community specific expression detection device according to the present invention.
  • FIG. 4 is a flowchart of the present invention.
  • FIG. 5 is a flowchart of document collection according to the present invention.
  • FIG. 6 is a flowchart for determining the suitability of an expanded word base.
  • CPU Central control unit

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Artificial Intelligence (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Computational Linguistics (AREA)
  • General Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Physics & Mathematics (AREA)
  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)

Abstract

 コミュニティ固有表現の収集に関係する従来技術では、専門的分野における名詞・複合名詞からなる専門用語の収集に関するものがあるが、名詞以外の新しい表現には応用が難しい。また、未知語・新語の収集の分野においても、対象はほぼ名詞に限定されて、新しい表現を規則的に収集するという手法は提案されていない。  所定のコミュニティで使用される文書集合の中から、(a)コミュニティに固有のnグラム連語を抽出する手段、(b)固有の表現の核となる可能性のある語基を選択する手段、(c)前記選択された語基をその前後に拡張する手段、(d)前記拡張された語基を文法に従って選別する手段によって、上記問題を解決している。

Description

コミュニティ特有表現検出装置及び方法
技術分野
[0001] 語形成論に基づき、コミュニティで使用される表現の中から当該コミュニティ特有の 表現を検出する装置及び方法に関する。
背景技術
[0002] 特定の興味やテーマをめぐって活発な議論が交されている人々のコミュニティにお いては、往々にしてそのコミュニティ独自の表現が発生する。例えば、 日本酒の味を 議論するコミュニティにおいては、「老ね(ヒネ)、ヒキのある、キレる、 · · ·」といった表 現が用いられる。ワインを好む人々の間では「フルボディ、ミディアムドライ、樽香、後 口、 · · ·」のような表現が見られる。これらは専門知識を有する人々の用いる難解な専 門用語ではなぐワインや日本酒の味に親しむ人であれば、その味を言い表す表現 として自然にその意味が理解される種類の語彙である。また、高校'大学生等の「若 者語」として集められている表現もコミュニティ固有の表現と考える事ができる。最近 では、インターネットの掲示板などに集まる人々の成すコミュニティにおいて多くの新 し!、表現が見 、だされるようになって!/、る。
特許文献 1:特開 2002-297589「未知語収集方法」
特許文献 2 :特開平 5-113997「辞書データ収集装置」
特許文献 3:特開 2004-265440「未知語登録装置および方法並びに記憶媒体」 特許文献 4 :特開 2005-309853「専門的記述と非専門的記述間の語彙変換方法 'プ ログラム'システム」
非特許文献 1 :中川祐志,湯本紘彰 , &辰則(2003).出現頻度と連接頻度に基づく専 門用語抽出. 自然言語処理, 10(1), 27-45.
非特許文献 2 :辻慶大, &芳鐘冬榭(2004).専門分野において重要となる新語の特 定に向けた基礎研究. 言語処理学会第 10回年次大会発表論文集, (pp. 189-191). 非特許文献 3 :藤井敦,伊藤克亘、秋葉友良(2003), IPA未踏ソフトウェア創造事業「 CYCLONE:最強事典サイトの構築」, www.ipa.go.jp/about/news/ event/pdf/29A7_f ujii.pdf
非特許文献 4:米川明彦 (1998)「若者語を科学する」東京:明治書院
発明の開示
発明が解決しょうとする課題
[0003] コミュニティ固有表現の収集に関係する既存技術には、主に専門用語の収集と未 知語の収集に関するものがある。専門用語の収集については、非特許文献 1、非特 許文献 2を始めとした研究があるが、ほとんどは専門的分野における名詞、複合名詞 力もなる専門用語の収集に関するものである。このように限定する事によって、単名 詞の重なりや連接関係等に着目したスコアに基づ 、たアルゴリズムを用いる事ができ る力 名詞以外の表現には応用が難しい。
また、未知語'新語の収集については、辞書の構築等においても重要なテーマで あり、特開 2002-297589「未知語収集方法」(特許文献 1)、特開 2004-265440「未知 語登録装置および方法並びに記憶媒体」(特許文献 3)等、既存特許にもこのテーマ を扱った技術が存在する。
[0004] し力しながら、非特許文献 3等の報告にもあるように日本語における未知語の検出 は困難な問題であり、特開 2002-297589「未知語収集方法」(特許文献 1)の方法もそ うである力 基本的には辞書に登録されていないものを人手ゃヒユーリステイクスによ つて収集しているものが多い。また、これら未知語の検出においても対象はほぼ名詞 に限定されており、真に新しい表現の収集という問題に焦点を絞ったものはまれであ る。
また、社会言語学において、高校生'大学生の用いる「若者語」の収集と分析を行う 分野が存在する(非特許文献 4)。コミュニティ固有の表現についての既存研究として は、本願発明に近いと思われるが、社会言語学分野で、若者語や流行語を規則的 に収集すると ヽぅ手法は提案されて ヽな ヽ。
課題を解決するための手段
[0005] 以下の装置を開示することにより課題を解決して!/、る。
(1)
以下の(a)から (d)の手段を有する所定のコミュニティで使用される文書集合力も前 記所定のコミュニティに特有な表現を検索する装置、
(a)前記コミュニティに特有に使用される nグラム連語を抽出する手段、
(b)前記特有な表現の核となる可能性のある第一の語基を選択する手段、
(c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択する手段
(d)前記拡張語基の中から当該言語の語形成規則に従って前記所定のコミュニティ に特有な表現を選別する手段。
[0006] (2)
さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデー タ検索することによって収集する手段を含むことを特徴とする (1)に記載の装置。 (3)
前記 nグラム連語を抽出する手段は、複数のコミュニティで使用される文書を用い、 前記所定のコミュニティで使用される nグラム連語の有意度と、他のコミュニティで使 用される nグラム連語との有意度との比較に基づいて前記 nグラム連語を抽出する手 段を含むことを特徴とする (1)及び (2)に記載の装置。
[0007] さらに、以下の方法を開示することにより課題を解決している。
(4)
以下の(a)から (d)のステップを有する、所定のコミュニティで使用される文書集合 から前記所定のコミュニティに特有な表現を検索する方法、
(a)前記コミュニティに特有に使用される nグラム連語を抽出するステップ、
(b)前記特有な表現の核となる可能性のある第一の語基を選択するステップ、
(c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択するステ ップ、
(d)前記拡張語基の中から当該言語の語形成規則に従って、前記所定のコミュ-テ ィに特有な表現を選別するステップ。 さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデー タ検索することによって収集するステップを含むことを特徴とする (4)に記載の方法。
[0008] さらに、以下のプログラムを開示することにより課題を解決している。
(6)
コンピュータを制御して、以下の(a)から (d)の手段を動作させ、所定のコミュニティ で使用される文書集合力 前記コミュニティに特有な表現を検索するプログラム、
(a)前記コミュニティに特有に使用される nグラム連語を抽出する手段、
(b)前記特有な表現の核となる可能性のある第一の語基を選択する手段、
(c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択する手段
(d)前記拡張語基の中から当該言語の語形成規則に従って前記所定のコミュニティ に特有な表現を選別する手段。
(7)
さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデータ 検索することによって収集する手段を含むことを特徴とする (6)に記載のプログラム。 発明の効果
[0009] 本願発明に従って、所望のコミュニティで使用される表現を収集しその意味を理解 することは、コミュニティのメンバーにとってコミュニケーションが容易になり、さらに、そ のアイデンティティを確認するのに役に立てることが出来る。また、そのコミュニティの 特徴や性格を分析する目的に役立てる事ができる。
さらに、商品の開発等においてユーザのコミュニティで交される議論の内容を分析 することが有用であると思われるが、この場合当該コミュニティ固有の表現を収集しそ の意味を理解する事は、この目的に大きく貢献すると考えられる。
また、本願発明は、主要品詞間の語法の拡張であり、他の言語にも応用可能であ る。英語の例を挙げると、「He 747'ed to Chicago.」という表現が可能である。これは 航空機の型番を動詞化したものである。また、「The web-logging is becoming a social phenomenon.」と!、う表現も用いられる力 これは「Web-log (ウェブに書き込む)」と!ヽ う動詞が名詞化された例である。
発明を実施するための最良の形態
[0010] 以下に最良の形態を説明する。
実施例 1
[0011] 図 1は、本願発明を実施する場合のシステム例を示している。ネットワーク 140には 、ユーザ PC110、サイトサーバ(1) 120、サイトサーバ(2) 130等が接続されている。 使用者がユーザ PC110を操作することにより、ネットワーク 140に接続されているサ イトサーバ(1) 120、サイトサーバ(2) 130等をアクセスし、検索ツール等を使用して 必要な情報を取得する。本願発明はインターネットでの検索を実施例として示すが、 これに限らず、情報が検索できるシステムならば他の方法でも応用できる。取得した 情報をユーザ PC上のコンピュータプログラムで処理し、所望の結果を得ることが出来 る。
[0012] 図 2は、本願発明の一部を実施するユーザ PCを示している。筐体 200の中には、 記憶装置 210、メインメモリー 220、出力装置 230、中央制御装置 (CPU) 240、操作 装置 250、ネットワーク 1/0260が含まれている。使用者が操作装置 250を操作し、ネ ットワーク I/Oを通して、必要な情報をインターネットの各サイトから入手する。中央制 御装置 240は記憶装置 210に記憶されている文書処理プログラムをメモリにダウン口 ードし、インターネットから検索された情報を用いて所定のデータ処理を行い出力装 置 230に結果を表示する。
[0013] 図 3は、本願発明によるコミュニティ固有表現検出装置のブロック図を示している。 3 10はコミュニティ文書検索部、 314はウェブサイト、 316は用語リスト格納部、 320は 文書処理部、 330は nグラム連語抽出部、 335は有意度判定部、 340は語基選択部 、 350は語基の左右拡張部、 354は左側拡張規則格納部、 356は右側拡張規則格 納部、 360は新表現の選別部、 365は言語規則格納部、 370は出力部を表す。 以下、これらの詳細について説明する。
[0014] [基本アルゴリズム]
図 4に示すフローチャートに従って、本願発明の基本アルゴリズムを説明する。
ステップ 410:コミュニティで使用される文書の収集 ステップ 420: nグラム連語の抽出
ステップ 430:新表現の核となる要素 (語基)の選択
ステップ 440:拡張語基の選択
ステップ 450:新 、表現の選別
[0015] [アルゴリズムの詳細]
以下にアルゴリズムの詳細について説明する。
(1)所定のコミュニティで使用される文書の収集(図 4 ステップ 410)
先ず、所定のコミュニティで使用される文書集合を次のステップで収集する。図 5に示 されるアルゴリズムを参照。
ステップ 510:用語の指定による候補文書の取得
ステップ 520:候補文書の前処理
ステップ 530:ノイズ文書の除去
ステップ 540:他のコミュニティ文書の検索の要否
以下、各ステップについて詳細に説明する。
[0016] (1— 1)ステップ 510 :候補文書の取得
本願発明を実施する為には、所定の用語を含む用語リストを用いて、所定のコミュ 二ティの関係者が使用する文書を収集する。ここで用語リストは用語リスト格納部(図 3 : 316)に格納されている。
ここで用語リストとは、一つのコミュニティにおけるキーワードとなる用語の集合であ る。例えば、一つのコミュニティとして「ワインの愛好家」を選択すると、用語リストの構 成要素は「ワインの銘柄」である。ワインの用語リスト中に記載されて 、る銘柄に従 、、 インターネットの検索ツールを使用して、ワインに関する情報を収集する(図 3 : 314) 。ここで、銘柄としては、 「ァウスレーゼ」、 「シャトー キユレ ボン」、 「シャトー マルゴ 一」、「ヴイン サント トスカーノ」等の銘柄を指定することが出来る。この用語をキー ワードとして、データベースから候補の文書を検索する。データベースとしてはこのよ うな情報が格納されて 、るデータベースならば何でも構わな 、が、本実施例ではイン ターネットの検索エンジンを使用して、候補の文書を検索する方法について説明する [0017] (1 2)ステップ 520 :候補文書の前処理
前処理では、先ずウェブページの情報力 文書に相当するものを取り出し文書解析 を行なう。次に、分かち書きを行ない内容語、助詞、助動詞等を抽出し、これらの文 書の特徴を表す特徴値を求める。これらの特徴値を用いて、以下の様にノイズ文書 を除去する。また、収集しょうとする文書の典型とみなされるような少量のモデル文書 を前もって選定しておく。
[0018] (1 3)ステップ 530 :ノイズ文書の除去
インターネットのウェブページから自動的にこれらの情報を収集した文書には様々 な情報が含まれており、そのままでは利用できない場合が多い。本実施例ではこれら の文書の中から、ガービッジ文書、リスト文書、及び日記型文書に該当する文書をノ ィズ文書として除去している。
以下に、ガービッジ文書、リスト文書、及び日記型文書について説明する。
(a)ガービッジ文書
内容語数が少ない文書、あるいは、固有名詞比率の低い文書等の条件の全てを満 足する文書を言う。内容語数とは、一つのウェブページに記載されている文書に含ま れているに内容語の数である。内容語とは助詞,助動詞を除いた、名詞、動詞、形容 詞、副詞に該当する単語である。また、ここで言う固有名詞とは、世間一般に固有名 詞であると認識されている名詞である。固有名詞比率とは一つのウェブページに出 現する固有名詞の数と内容語数との比率である。
(b)リスト文書
固有名詞比率が高い文書、内容語と助詞'助動詞との相関係数が低い文書等の条 件の全てを満足する文書をリスト情報文書と定義する。これはインターネットのサイト において、ある領域における対象物に関する情報が単なるリストとして格納されてい る文書である。
[0019] (c)日記型文書
あるコミュニティに関する固有名詞比率が低 、文書、内容語 nグラムに基づくモデル 文書との相関度が低い文書、助詞'助動詞 nグラムに基づくモデル文書との相関度が 高い文書等の条件の全てを満足する文書を日記型文書と定義する。これらは言わば 個人の日記書き込みサイトとして利用されている文書、及び、デパートの売場に関す るサイトなど、主として他の情報が記載されて要る文書である。以上の定義に基づい て、ガービッジ文書、リスト文書、及び、日記型文書をノイズ文書として除去する。
[0020] (1 -4)ステップ 540:他のコミュニティ文書の検索の要否
ステップ 510からステップ 530により、所定のコミュニティで使用される文書集合が収 集される。ステップ 540では、他のコミュニティで使用される文書集合を同様に収集す る。
[0021] 次にこれらの収集された複数のコミュニティで使用される文書集合を用いて、これら のコミュニティで固有に使用される新しい表現を選別する。
以上により、複数のコミュニティで使用される文書集合が作成される(図 3: 320)。
[0022] (2) nグラム連語の抽出(図 4 ステップ 420)
(2— 1)コミュニティ固有の連語抽出
単語レベルの n-gram連語 (nグラム連語)を統計的手法により、特定のコミュニティで 使用される場合に有意に出現するものを抽出する。これらをコミュニティ固有の連語と 呼ぶ。これらの詳細について説明する。
nグラム連語とは、連続した 1以上の語であって、一語の場合はュ-グラム(Uni-gra m)、二語の場合はバイグラム (Bト gram)、三語の場合はトライグラム (Tri-gram)と呼 ばれる。本実施例では、バイグラム、トライグラムを用いている(図 3 : 330)。
[0023] (2— 2)有意度による判定
単純に nグラム連語を求めると数多くの nグラム連語が得られる力 全ての nグラム連 語が有効であるとは限らない。そこで、二つのコミュニティで使用される文書集合を比 較し、一方のコミュニティで使用されている nグラム連語が、一方に有意に偏って出現 する nグラム連語を選択する (Z検定)。本願明細書では、 2つの文書集合においてそ れぞれの nグラム連語の出現する比率を比較し、その比率差を検定する方法を用い る(図 3 : 330)。 ここで、ある nグラム連語 Wが 2つの文書集合 dl, d2に共に表れたと 考え、その頻度力 Swl, w2であったとする。文書集合 dlに表れた用語の総数を nl,文 書 d2のそれを n2とする。すると Wがそれぞれの文書集合に表れた割合は次のように なる。 [0024] (式 l) pl= wl/nl,
(式 2) p2= w2/n2
ここで、標本比率を実際のデータから得られた比率とすると、 pi及び p2は標本比率で ある。
ここで、 pi > p2である場合に、これが有意であるかどうかを検定する、すなわち、 nグ ラム連語 Wは dlの文書の方に有意に偏って出現するかどうかを検定するということを 意味する (片側検定)。
ここで、帰無仮説と対立仮説は次のようになる。
HO: pil = pi2 帰無仮説
HI: pil > pi2 片側検定における対立仮説
検定を行うために、まず実際には知られて ヽな 、母比率 pihat (式 3)を標本比率から 推定する。
(式 3) pihat = (nl*pl + n2*p2) I (nl + n2)
ここから zを (式 4)で計算する。
(式 4) z = (pl-p2)/ (pihat*(l- pihat)*(l/nl+l/n2》
帰無仮説を棄却し、対立仮説を採用するには、 5%の危険率において、 z > 1.65で なくてはならないことになる。
[0025] このようにして、全ての連語にっ 、て検定を行 、、文書集合の中に現れる nグラム 連語であって、一方のコミュニティで使用される文書に有意に出現する nグラム連語、 及び、他方のコミュニティで使用される文書に有意に出現する nグラム連語をそれぞ れ選択することが出来る。従って、双方のコミュニティで共通に使用されるものは選択 されないこととなる。
本願実施例では、ワインの愛好家が使用する文書集合と、 日本酒の愛好家が使用 する文書集合とに特徴的に現れる 2グラム、 3グラムのリストを取り出し、 Z検定を行って いる。ここで、 Z検定の結果、ワインの愛好家が使用する文書集合から、 Z値が 1.65以 上の nグラムを選択する。
[0026] (3)新表現の核となる要素 (語基)の選択(図 4 ステップ 430)
ここで、上記の方法によって抽出された nグラムについて、その中力 新しい表現の 核となる要素を取り出す(図 3 : 340)。そのためには、 nグラム連鎖をひとまず切り離し 、そこに生じる全ての要素(形態素)のリストを作る。そこから、核となる可能性のない ものを除外する。ここで、核となる可能性の無いものとしては、助詞、助動詞、接続詞 、活用語尾等の機能語、「、」、「。」、「?」等の切れ目要素がある。また、「平仮名一 文字」、「片仮名一文字」のものも除外する。これによつて、新表現の核となる可能性 のある要素のリスト (核リスト)が作成される。
[0027] (4)拡張語基の選択(図 4 ステップ 440)
(4 1)語基の拡張
それぞれの語基候補について、連語パターンの分布に基づき、前後の要素を取り 込んで拡張する必要があるかどうかを判断する(図 3: 350)。
ここで、(式 5)の様に Z を定義する。
ratio
(式 5) Z = Z[X]/AvgZ([X][X+l])
ratio
ここで、 Z[X]とは現在着目している nグラム語基の Z値である。核要素を Xとし、それに 1語拡張した要素を [X + 1]とし、 2語拡張した要素を [X+2]とする。 AvgZ([X][X+l])と は nグラム語基から「右」に一語拡張したときの、 [X][X+1]に相当する全ての(n+ 1)グ ラムの語基の Z値の平均値である (0<Z ;)。
ratio
正確に言えば、 nグラム語基から「左」に一語拡張したときの AvgZ([X- 1][X])も考えら れる。従って、以下本願明細書では、 Z と言ったときには、特段の説明がない限り、
ratio
nグラム語基から「左」または「右」に一語拡張したときの双方を含むものとする。さらに 、データ処理の便宜上、 Z の対数をとつて、(式 6)を定義する。
ratio
(式 6) LZ = 10 * log(Z )
ratio
[0028] (4 2)右側拡張規則
図 6のアルゴリズムに示すように、 nグラム語基力 右に一語拡張したときには、以下 の規則を適用する(図 3 : 356)。但し、 [X+l]、及び、 [X+2]の最後の語が切れ目要素 の場合には除外する。
[0029] 第一の条件
(i) Z([X],[X+l]) > Avg Ζ([Χ],[Χ+1],[Χ+2])、かつ、
(ii) LZ > 第 1閾値 を満たす場合には [X+1]へ拡張する候補として選択される(610、 620、 650)。ここで、 第 1の閾値は本実施例では 5.0とし、 Z([X],[X+1])は ([X],[X+1])で表現される (n+ 1) グラム語基の Z値、 AvgZ([X],[X+l],[ X+2])は、 [X],[X+1],[ X+2]に相当する全ての(n + 2)グラムの Z値の平均値である。なお、第一の条件で使用される LZに対する第一 の閾値は高く設定されている。この値が高い場合には、 Zの値による判定のみでも十 分に新表現として認定され得ると判断されるので、 Jratio (後述)の値に関わらず、新 表現の可能性のあるものとして選択する。
第一の条件、すなわち、(i)及び (ii)の双方の条件を満足する場合には、拡張された 語基の候補として選択される (650)。(i)の条件を満たさな!/、場合には拡張する候補 としては選択されない (660)。(i)の条件を満たすが、(ii)の条件を満たさない場合に は、次に示す第二の条件で判別する(630、 640)。
[0030] 第二の条件
(iii) LZ >第 2閾値、かつ、
(iv) Jratio = Njun/Nall > 第 3閾値
を満たす場合には [X+1]へ拡張する候補として選択される(630、 640、 650)。
第二の条件で使用される LZに対する第二の閾値は、実施例では 3.0と設定されて おり、 LZがこの値より大きぐかつ、 Jratioが 0.1以上の値を取る時にのみ、新表現の 可能性があると判定される。
ここで、 Jratioとは [X+2]要素が切れ目要素である割合のことである(0=く Jratio =< D oまた、第 3閾値は本実施例では 0.1とし、 Njunは切れ目要素と認定された先端要 素 [X+2]の数、 Nailは対象となった [X+2]に相当する (n+2)グラムの数である。
第二の条件、すなわち、(iii)及び (iv)の双方の条件を満足する場合には、拡張さ れた語基の候補として選択される (650)。(iii)及び (iv)の 、づれかの条件を満たさな V、場合には拡張された語集は選択されな 、 (660)。
[0031] (4 3)左側拡張規則
基本的に右側拡張規則と同様である(図 3 : 354)。前述の (i)、 (ii) , (iii)の条件は 全く同じである。但し、 (iv)において、切れ目要素のカウント方法が異なる。右側拡張 規則では、 [老] [ねる]のような例に現れる [ねる]のように、着目している動詞の活用語 尾は切れ目要素と見なさない。しかし、左側拡張規則では、着目している語基の左側 に存在する動詞の活用語尾が、着目して ヽる語基の新し ヽ表現の接頭辞として用い られることは考えにくい。従ってこの場合には切れ目要素としてカウントされる。すな わち、左側では切れ目要素としてカウントされる要素が追加される。
[0032] (4 4)右側拡張規則適用例
右側拡張規則について実例を使って説明する。語基として選択されたフルーティー (Z値は 147.14)を右側に拡張することについて説明する。
ロロ基 拡張 Z値
[X] [X+1] [X+2]
[フル -ティ -] [さ] 5.66
[フル -ティ -] [さ] [が] 2.00
[フル —ティ一] [さ] [は] 2.00
ここで、注目している語基は、「フルーティー」である。先ず、右に一個延ばして検討 する。 [フルーティー]、 [さ]は前述の [X] [X+ 1]に対応する。
[0033] この時の Z値は以下のようになる。
Z([X][X+ 1]) =Z ([フルーティー] [さ]) = 5.66
さらに右に一個延ばして ([X][X+ 1][Χ+ 2])を検討する。ここでは 2つの連語が見つ かる。すなわち [フルーティー] [さ] [が]、及び、 [フルーティー] [さ] [は]である。
[フルーティー] [さ] [が]の Z値 =Z ([フルーティー] [さ] [が]) =2.00
[フルーティー] [さ] [は]の Z値 =Z ([フルーティー] [さ] [は]) =2.00
ここで、 [X+ 2]の要素、すなわち、「が」「は」を kOne要素と呼ぶ。この例のように複数 の kOne要素がある場合には、これらの Z値の平均値を求める。この場合、どちらも 2.0 0であるので、平均値は 2.00となる。
すなわち、 AvgZ([X][X+l][X+2]) = 2.00、次に LZを求める。
Zratio = Z([X] [X + 1] ) / AvgZ([X] [X+ 1] [X+2]) = 5.66/2.00 = 2.83
LZ= 10*log(Zratio)= 4.52となる。
[0034] 次に、この kOne要素について、切れ目を示す「切れ目要素」であるかどうかを調べ る。すなわち, 「フルーティーさ」という新しい表現の候補の後に、文法的切れ目を示 す要素があるかどうかをチェックする。もしあれば、その候補(「フルーティーさ」)が文 法的にひとまとまりの要素として扱われていることを示唆し、新表現の候補となる。ここ では、「が」「は」共に格助詞であり、文法的切れ目を示す要素である。つまり要素(「 フルーティーさ」)とつながってさらに大きな一まとまりの表現や語を作ることは考えに くい。 kOne要素のうち切れ目要素である割合を Jratioと呼ぶ。ここでは、 2つとも切れ 目要素であるから、 Jratio = 2/2 = 1となる。
[0035] これらの準備をした上で、新表現としての可能性のあるものを検出していく。先ず、 第一の条件について検討する。
第一の条件
(i) Z([X],[X+l]) >AvgZ([X],[X+l],[X+2])、かつ、
(ii) LZ>第 1閾値
(i)の条件は、 Z ([フルーティー] [さ]) =5.66、及び、 AvgZ([X][X+l][X+2]) = 2.00であ るので満足する。
(ii)の条件は、 LZ= 10*log(Zratio)= 4.52、第 1閾値 =5.0となり、この条件を満足しな い。従って第一の条件は満足しないので、次に第二の条件について検討する。
[0036] 第二の条件
(iii) LZ >第 2閾値、かつ、
(iv) Jratio = NjunZNall>第 3の閾値
(iii)の条件は、 LZ=4.52、第 2の閾値は 3.0であるので満足する。(iv)の条件は、 Jra tio = 2/2 = 1、であり、第 3の閾値は 0.1であるので、満足する。
以上より、第二の条件を満足するので、 [フルーティー]から [フルーティーさ]へ拡張 される。ちなみに [フルーティーさ]の Z値 =Z ([フルーティー] [さ]) =5.66である。
[0037] (4 5)左側拡張規則適用例
左側拡張規則について実例を使って説明する。語基として選択された [受け] (Z値 は 73.01)を左側に拡張することについて説明する。
語基 拡張 Z値
[X-2] [X-1] [X]
[も] [受け] 6.83 [に] [も] [受け] 2.83
[女性] [受け] 6.83
[女性] [受け] 2.00
[あまり] [女性] [受け] 2.00
右側拡張規則の例と同様であるので左側にも拡張する。
[0038] 先ず、第一の条件について検討する。
(i) Z([X-l],[X]) > Avg Z([X],[X-1],[ X-2])、かつ、
(ii) LZ >第 1閾値
Z([X-1][X]) = 6.83、及び、 AvgZ([X][X- 1][X- 2] = 2.00であるので、(i)の条件は満足 する。 LZ=5.33、第 1閾値は 5.0であるので、(ii)の条件も満足する。
以上より、 [受け]から [女性受け]へ拡張される。ちなみに [女性受け]の Z値 =Z ([女 性受け]) =5.33である。
[0039] (5)新しい表現の選別(図 4 ステップ 450)
拡張の条件に合致するものの中から、語形成規則に合致するものを新表現として 選び出す(図 3 : 360)。新しい表現を生み出す可能性の高い語は、日本語形成規則 に従っていなくてはならず、その形成規則は限られている(図 3 : 365)。新しい表現と して選別するためには、語法の拡張の起きている部分が名詞、動詞、形容詞、形容 動詞等を形成するための規則を遵守しているか否かを確認する必要がある。図 7に 示すフローチャートに従って説明する。
710 :名詞化規則
720 :動詞化規則
730 :形容詞化規則
740 :形容動詞化規則
750:全ての条件を満たさな!/、場合は候補として選別しな!、。
760:何れかの条件を満たす場合は候補として選別する。
以下詳細に説明する。
[0040] (5— 1)名詞化規則 (ステップ 710)
名詞化形成規則に合致するものは、語基の拡張の候補として選別される。名詞化と しては、「語基 +接尾辞」、「動詞連用形名詞化」、「複合名詞」などが挙げられる。そ れぞれにつ 、て、日本語としての規則を満足して ヽるカゝ確認する必要がある。
(a)語基 +接尾辞
名詞以外の形容詞などを名詞化する場合は、それらの語尾に「さ」、「み」などを追カロ する場合がある。例として以下のものが挙げられる。
「さ」 (薄さ、悲しさ、ほめられたさ)
「け」 け、ねむけ、吐さけ、力ざりけ)
「み」 (強み、いやみ、すごみ)
[0041] (b)動詞連用形名詞化
語基の右側に格助詞'名詞をつけることによって、動詞連用形を名詞用法する場合 も起こり得る。例えば、以下の様な例が挙げられる。
「走る」から「走り」、「歩き」
「遊ぶ」から「遊び」
(c)複合名詞
複合名詞とみなされるものは、語基の拡張の候補として選別される。例えば、以下の 様な例が挙げられる。
語尾に [米]を付けた場合 [掛け] [米]、 [麹] [米]、 [純] [米]、 [赤] [米] 語尾に [香]を付けた場合 [バナナ] [香]、 [吟醸] [香]、 成] [香]
(d)英語の名詞化について
本願発明は日本語だけでなく外国語にも応用できる。英語を例にとって説明する。 英語で元来名詞以外の品詞として使用されて ヽるものが名詞として使用されて ヽる 場合がある。例えば、以下の様な接尾辞を付加することによって名詞化されている。 「ness」: pleasantness, ugliness
「ing」: gatnermg
「ful」: earful
「dom」: femidom
「hood」: broherhood, womanhood
[0042] (5— 2)動詞化規則 (ステップ 720) 動詞化形成規則に合致するものも、語基の拡張の候補として選別される。動詞化の 例として「名詞 +する」、「動詞の一般活用形」等が考えられる。拡張の候補として選 択されたものが、 日本語としての規則を満足して ヽるか確認する必要がある。
(a)「名詞 +動詞化接尾辞」の形態であるか
名詞に「する」、「ぶる」のような動詞化接尾辞、またはその活用形が結合されている 場合には、語基の動詞化拡張の候補として選別される。例えば、「お茶」に「する」を つけて「お茶する」とする場合, 「美人」に「ぶる」をつけて「美人ぶる」が挙げられる。
(b)動詞の一般活用形
拡張された語基が、「名詞 +動詞化接尾辞」の形態を除いた、動詞の一般活用形 である場合にも語基の拡張の候補として選別される。例えば,名詞に動詞の活用語 尾をつけて動詞化してしまう生産的な例として、以下のような例が挙げられる、「デモ る、デモらない,デモれば」。同様に, 「ゲバる、ハモる、ツモる、ダーグる」といった新 L ヽ動詞をこの方法で作ることができる。
[0043] (c)英語の動詞化について
本願発明は日本語だけでなく外国語にも応用できる。英語を例にとって説明する。 英語で元来名詞として使用されて ヽるものが動詞として使用されて ヽる場合がある。 Are you googling?
元来名詞である「google」が「googleを使って検索する」という動詞として使用されてい る例である。
I 747 ed to Chicago.
元来航空機の型番である「747」が「747航空機に乗った」 t 、う動詞として使用され ている例である。
その他、以下の様な接尾辞によって動詞化されている。
| ify」: Frenchify
「en」: enliven, soften
I izej: pluralize
[0044] (5— 3)形容詞化規則 (ステップ 730)
形容詞化形成規則に合致するものも語基の拡張の候補として選別される。拡張の 候補として選択されたものが日本語としての規則を満足して ヽるか確認する必要があ る。
「い」(しんどい、四角い)
「こい」(ネチつこい)
「ぼい」(おんなっぽい、それっぽい)
[0045] (5— 4)形容動詞化規則 (ステップ 740)
形容動詞化形成規則に合致するものも語基の拡張の候補として選別される。拡張 の候補として選択されたものが日本語としての規則を満足して 、る力確認する必要が ある。
「風」(王朝風、レゲ一風)
「な」(マックな [人])
「げ」(うれしげ、よさげ、なにげ)
以上のステップ 710からステップ 740までの何れかの条件を満足する場合には、語 基の拡張の候補として選別される(760)。いずれの条件も満たさない場合には、語基 の拡張の候補として選別されな ヽ (750)。
[0046] [実験結果]
以上のアルゴリズムに従って、実際のデータを用いた実験結果を示す。なお、本実 験では、対象とするコミュニティとして「日本酒の味覚を議論するコミュニティ」と「ワイ ンの味覚を議論するコミュニティ」を例として取り上げている。 日本酒およびワインの 銘柄名を「キーワード」として、インターネットの検索ツールを使用してそれぞれの文 書集合を収集した。
[0047] (1)名詞化
(1 1)語基 +接尾辞
形容詞を名詞化する例について説明する。ここでは形容詞「フルーティ」を名詞化 し「フルーティさ」とする例について説明する。
語基 拡張 Z値
[X] [X+1] [X+2]
[フルーティー] [さ] 5.66 [フルーティー] [さ] [が] 2.00
[フルーティー] [さ] [は] 2.00
[フルーティー]から [フルーティーさ]へ拡張されることは前述のとおりである。
次に、拡張された語基が名詞化形成規則 (語基 +接尾辞)を満足するか否か検討 する。名詞以外の形容詞などを名詞化する場合は、これらの語に「さ」、「み」などを追 加する。この実施例ではこの条件を満足している。
以上より、新 、語基として「フルーティー」の名詞である「フルーティーさ」が選択さ れる。ちなみに、「フルーティー」 +「さ」の判定のための LZ値は 4.52である。
(1 2)動詞連用形名詞化
語基として選択された [受け] (Z値は 73.01)を左側に拡張することについて説明する 拡張 ロロ z値
[X- 2] [X- 1] [X]
[も] [受け] 6.83
[に] [も] [受け] 2.83
[女性] [受け] 6.83
[ゝ ] [女性] [受け] 2.00
[あまり] [女性] [受け] 2.00
[受け]から [女性受け]へ拡張されることは前述の通りである。そこで、拡張された語 基が規則 (動詞連用形名詞化)を満たすか否か検討する。 [女性]は名詞であることは 明らかである。また [受け]は後ろに格助詞が続く連語が見られ、動詞連用形による名 詞化がなされていると考えられることから、 [女性] [受け]は動詞連用形による名詞化で あると考えられるのでこの条件も満足する。
以上より、新しい語基として [女性] [受け]が選択される。ちなみに、 [女性] [受け]の判 定のための LZ値は 5.33である。
(1 3)複合名詞
語基として選択された [雪] (Z値は 66.96)を左側に拡張することについて説明する。
語基 拡張 Z値 [X] [X+1] [X+2]
園 [の] 4.00
園 [の] [中] 2.00
園 [温] 4.00
園 [で] 2.00
[雪] [室] 4.00
前述の条件にあてはめて検討すると [雪]から [雪温]へ拡張されることが分かる。ここ での詳細な説明は割愛する。次に拡張された語基が名詞化形成規則 (複合名詞)を 満足するか否か検討する。 [雪]及び [温]は名詞であることは明らかであるのでこの条 件も満足する。
以上より、新しい語基として [雪温]が選別される。ちなみに、 [雪温]の判定のための
LZ値は 3.01である。
その他の複合名詞として拡張された例としては以下のものがある。
[米]を語基として、 [掛け] [米]、 [麹] [米]、圆 [米]、 [赤] [米]
[香]を語基として、 [バナナ] [香]、 [吟醸] [香]、 成] [香]
[様]を語基として、 [マスカット] [様]、 [リンゴ] [様]、 [果実] [様]
[度]を語基として、 [アミノ酸] [度]、 [アルコール] [度]、 本酒] [度]
(2)動詞化
(2— 1)「名詞 +動詞化接尾辞」
「名詞 +する」の様な動詞化パターンの検出について説明する。ここでは、語基として 「悪酔!ヽ」 (Z値は 24.01である)を選択し右側へ拡張する。
左側拡張 語基 Z値
[X-2] [X-1] [X]
[悪酔い] [する] 4.00
[から] [悪酔い] [する] 2.00
[使用] [する] 2.00
前述の条件にあてはめて検討すると「悪酔 ヽ」を「悪酔 、する」へ拡張し新 、語基 とすることが出来る。ここでの詳細な説明は割愛する。 [0051] 次に、拡張された語基が動詞化規則(「名詞 +する」)を満足する力否かについて 検討する。この例では、名詞に「する」または「する」の活用形が結合されているので、 この条件を満たす。
以上より、新しい語基として「悪酔いする」が選別される。ちなみに、 [雪温]の判定の ための LZ値は 3.01である。
ここで、「悪酔いする」は普通に使用される言葉であると考えられる力 「ワインの味 覚を議論するコミュニティ」と比較して、「日本酒の味覚を議論するコミュニティ」では 有意差を持って出現していることが分力る。
その他の動詞化として拡張された例としては以下のものがある。
[醸造]を語基として [醸造] [する]、 [調和]を語基として [調和] [する]、 [登場]を語基とし て [登場] [する]、 [倍増]を語基として [倍増] [する]
[0052] (2— 2)動詞の一般活用形
動詞が文法に従って活用する場合に、「語基 +拡張部」がー個の新しい動詞を形 成する例について説明する。
例えば、日本酒コミュニティで用いられるパターンから、 [老] [ね] (読み:ひね)、 [老] [ ねた] (読み:ひねた)、 [老] [ね] [が、を (格助詞)] (読み:ひねが、ひねを)等のデータ が得られる。
語基 右側拡張 Z値
[老] [ねる] (読み:ひねる) 2.05
[老] [ねた] (読み:ひねた) 2.05
前述のアルゴリズムに従って、老ねる(読み:ひねる)(動詞一段活用形)が候補として 選択される。ここで、 [老] (読み:おい)は、一般名詞として辞書に登録されており、動 詞としては [老いる] (読み:おいる)という上一段動詞が登録されている。データと動 詞活用規則から、 [老ねる] (読み:ひねる)という下一段動詞としての拡張が起きてい ると判断される。また、 [老] [ね] + [格助詞]等のデータから、動詞連用形 [老ね] (読み: ひね)が名詞として用いられる名詞化が起きていることが分かる。ここから、 [老ねる] ( 読み:ひねる)がこのコミュニティにお 、て新 、表現として共通の言葉として使用さ れている事が推測される。 図面の簡単な説明
[0053] [図 1]本願発明を実施するシステム例を示す図である。
[図 2]本願発明の一部を実施する PCのブロック図である。
[図 3]本願発明によるコミュニティ固有表現検出装置のブロック図である。
[図 4]本願発明のフローチャートである。
[図 5]本願発明の文書収集のフローチャートである。
[図 6]拡張した語基の適否を判断するフローチャートである。
[図 7]拡張した語基が語形成規則に合致しているかを判定するフローチヤ 符号の説明
[0054] 110:ユーザ PC
120:サイトサーバ(1)
130:サイトサーバ(2)
140:ネットワーク
200:筐体
210:記憶装置
220:メインメモリー
230:出力装置
240:中央制御装置(CPU)
250:操作装置
260:ネットワーク I/O

Claims

請求の範囲
[1] 以下の(a)から (d)の手段を有する、所定のコミュニティで使用される文書集合から 前記所定のコミュニティに特有な表現を検索する装置、
(a)前記コミュニティに特有に使用される nグラム連語を抽出する手段、
(b)前記特有な表現の核となる可能性のある第一の語基を選択する手段、
(c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択する手段
(d)前記拡張語基の中から当該言語の語形成規則に従って前記所定のコミュニティ に特有な表現を選別する手段。
[2] さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデー タ検索することによって収集する手段を含むことを特徴とする請求項 1に記載の装置
[3] 前記 nグラム連語を抽出する手段は、複数のコミュニティで使用される文書を用い、 前記所定のコミュニティで使用される nグラム連語の有意度と、他のコミュニティで使 用される nグラム連語との有意度との比較に基づいて前記 nグラム連語を抽出する手 段を含むことを特徴とする請求項 1及び 2に記載の装置。
[4] 前記拡張語基を選択する手段は、さらに、
前記第二の語基の数、及び、前記第二の語基に取込まれた要素が切れ目要素であ る数を用いて算出された値に基づいて前記拡張語基を選択する手段含むことを特徴 とする請求項 1及び 2に記載の装置。
[5] 前記語形成規則に従って選別する手段は、名詞化規則、動詞化規則、形容詞化 規則、及び、形容動詞化規則のうち少なくとも 1つの語形成規則を含むことを特徴と する請求項 1及び 2に記載の装置。
[6] 以下の(a)から (d)のステップを有する、所定のコミュニティで使用される文書集合 から前記所定のコミュニティに特有な表現を検索する方法、
(a)前記コミュニティに特有に使用される nグラム連語を抽出するステップ、
(b)前記特有な表現の核となる可能性のある第一の語基を選択するステップ、 (c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択するステ ップ、
(d)前記拡張語基の中から当該言語の語形成規則に従って、前記所定のコミュ-テ ィに特有な表現を選別するステップ。
[7] さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデー タ検索することによって収集するステップを含むことを特徴とする請求項 6に記載の方 法。
[8] 前記 nグラム連語を抽出するステップは、複数のコミュニティで使用される文書を用 い、前記所定のコミュニティで使用される nグラム連語の有意度と、他のコミュニティで 使用される nグラム連語との有意度との比較に基づいて前記 nグラム連語を抽出する ステップを含むことを特徴とする請求項 6及び 7に記載の方法。
[9] コンピュータを制御して、以下の(a)から (d)の手段を動作させ、所定のコミュニティ で使用される文書集合力 前記コミュニティに特有な表現を検索するプログラム、
(a)前記コミュニティに特有に使用される nグラム連語を抽出する手段、
(b)前記特有な表現の核となる可能性のある第一の語基を選択する手段、
(c)前記第一の語基の有意度、及び、前記第一の語基の前又は後の要素を取込ん だ第二の語基の有意度を用いて算出された値に基づいて拡張語基を選択する手段
(d)前記拡張語基の中から当該言語の語形成規則に従って前記所定のコミュニティ に特有な表現を選別する手段。
[10] さらに、前記文書集合を、所定の用語リストに含まれる用語をキーワードとしてデー タ検索することによって収集する手段を含むことを特徴とする請求項 9に記載のプロ グラム。
[11] 前記 nグラム連語を抽出する手段は、複数のコミュニティで使用される文書を用い、 前記所定のコミュニティで使用される nグラム連語の有意度と、他のコミュニティで使 用される nグラム連語との有意度との比較に基づいて前記 nグラム連語を抽出する手 段を含むことを特徴とする請求項 9及び 10に記載のプログラム。
PCT/JP2006/314000 2005-07-15 2006-07-13 コミュニティ特有表現検出装置及び方法 Ceased WO2007010836A1 (ja)

Priority Applications (4)

Application Number Priority Date Filing Date Title
CN2006800258021A CN101223521B (zh) 2005-07-15 2006-07-13 社群特有表现检测装置及方法
JP2007525983A JPWO2007010836A1 (ja) 2005-07-15 2006-07-13 コミュニティ特有表現検出装置及び方法
US11/990,495 US20100076745A1 (en) 2005-07-15 2006-07-13 Apparatus and Method of Detecting Community-Specific Expression
DE112006001822T DE112006001822T5 (de) 2005-07-15 2006-07-13 Vorrichtung und Verfahren zum Erfassen eines gemeinschaftsspezifischen Ausdrucks

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
JP2005207810 2005-07-15
JP2005-207810 2005-07-15

Publications (1)

Publication Number Publication Date
WO2007010836A1 true WO2007010836A1 (ja) 2007-01-25

Family

ID=37668717

Family Applications (1)

Application Number Title Priority Date Filing Date
PCT/JP2006/314000 Ceased WO2007010836A1 (ja) 2005-07-15 2006-07-13 コミュニティ特有表現検出装置及び方法

Country Status (6)

Country Link
US (1) US20100076745A1 (ja)
JP (1) JPWO2007010836A1 (ja)
KR (1) KR20080024530A (ja)
CN (1) CN101223521B (ja)
DE (1) DE112006001822T5 (ja)
WO (1) WO2007010836A1 (ja)

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010160534A (ja) * 2009-01-06 2010-07-22 Yahoo Japan Corp 地域特性辞書生成方法及び装置
JP7557770B2 (ja) 2020-06-05 2024-09-30 国立大学法人北海道国立大学機構 専門用語抽出装置、専門用語抽出方法及びプログラム

Families Citing this family (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US8473279B2 (en) * 2008-05-30 2013-06-25 Eiman Al-Shammari Lemmatizing, stemming, and query expansion method and system
US8423350B1 (en) * 2009-05-21 2013-04-16 Google Inc. Segmenting text for searching
US20110082687A1 (en) * 2009-10-05 2011-04-07 Marcelo Pham Method and system for taking actions based on analysis of enterprise communication messages
KR101706827B1 (ko) * 2014-12-04 2017-02-16 강원대학교산학협력단 개체 간 사회 관계 추출 장치 및 방법

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH1185761A (ja) * 1997-09-03 1999-03-30 Ee I Soft Kk 未知語登録装置および方法並びに記録媒体
JP2004062262A (ja) * 2002-07-25 2004-02-26 Hitachi Ltd 未知語を自動的に辞書へ登録する方法

Family Cites Families (22)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US5265065A (en) * 1991-10-08 1993-11-23 West Publishing Company Method and apparatus for information retrieval from a database by replacing domain specific stemmed phases in a natural language to create a search query
US5799268A (en) * 1994-09-28 1998-08-25 Apple Computer, Inc. Method for extracting knowledge from online documentation and creating a glossary, index, help database or the like
US5704060A (en) * 1995-05-22 1997-12-30 Del Monte; Michael G. Text storage and retrieval system and method
US6173298B1 (en) * 1996-09-17 2001-01-09 Asap, Ltd. Method and apparatus for implementing a dynamic collocation dictionary
US5933822A (en) * 1997-07-22 1999-08-03 Microsoft Corporation Apparatus and methods for an information retrieval system that employs natural language processing of search results to improve overall precision
GB2338089A (en) * 1998-06-02 1999-12-08 Sharp Kk Indexing method
US6347316B1 (en) * 1998-12-14 2002-02-12 International Business Machines Corporation National language proxy file save and incremental cache translation option for world wide web documents
US6442524B1 (en) * 1999-01-29 2002-08-27 Sony Corporation Analyzing inflectional morphology in a spoken language translation system
US6356865B1 (en) * 1999-01-29 2002-03-12 Sony Corporation Method and apparatus for performing spoken language translation
US7865358B2 (en) * 2000-06-26 2011-01-04 Oracle International Corporation Multi-user functionality for converting data from a first form to a second form
US7225199B1 (en) * 2000-06-26 2007-05-29 Silver Creek Systems, Inc. Normalizing and classifying locale-specific information
US8396859B2 (en) * 2000-06-26 2013-03-12 Oracle International Corporation Subject matter context search engine
US6675159B1 (en) * 2000-07-27 2004-01-06 Science Applic Int Corp Concept-based search and retrieval system
US7526425B2 (en) * 2001-08-14 2009-04-28 Evri Inc. Method and system for extending keyword searching to syntactically and semantically annotated data
WO2005024604A2 (en) * 2003-09-09 2005-03-17 Siftology, Inc. Dynamic lexicon
US20050149510A1 (en) * 2004-01-07 2005-07-07 Uri Shafrir Concept mining and concept discovery-semantic search tool for large digital databases
US7260568B2 (en) * 2004-04-15 2007-08-21 Microsoft Corporation Verifying relevance between keywords and web site contents
US20070217693A1 (en) * 2004-07-02 2007-09-20 Texttech, Llc Automated evaluation systems & methods
US7571157B2 (en) * 2004-12-29 2009-08-04 Aol Llc Filtering search results
WO2006096260A2 (en) * 2005-01-31 2006-09-14 Musgrove Technology Enterprises, Llc System and method for generating an interlinked taxonomy structure
US7657421B2 (en) * 2006-06-28 2010-02-02 International Business Machines Corporation System and method for identifying and defining idioms
US7698328B2 (en) * 2006-08-11 2010-04-13 Apple Inc. User-directed search refinement

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JPH1185761A (ja) * 1997-09-03 1999-03-30 Ee I Soft Kk 未知語登録装置および方法並びに記録媒体
JP2004062262A (ja) * 2002-07-25 2004-02-26 Hitachi Ltd 未知語を自動的に辞書へ登録する方法

Non-Patent Citations (2)

* Cited by examiner, † Cited by third party
Title
MORI S. ET AL.: "n Glam Tokei ni yoru Corpus kara no Michigo Chushutsu", IEICE TECHNICAL REPORT NLC 95-8, vol. 95, no. 168, 20 July 1995 (1995-07-20), pages 7 - 12, XP003007716 *
NAGAO M. ET AL.: "Daikibo Nihongo Text no n glam Tokei no Tsukurikata to Goku no Jido Chushutsu 93-NL-96-1", vol. 93, no. 61, 9 July 1993 (1993-07-09), pages 1 - 8, XP003007717 *

Cited By (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
JP2010160534A (ja) * 2009-01-06 2010-07-22 Yahoo Japan Corp 地域特性辞書生成方法及び装置
JP7557770B2 (ja) 2020-06-05 2024-09-30 国立大学法人北海道国立大学機構 専門用語抽出装置、専門用語抽出方法及びプログラム

Also Published As

Publication number Publication date
KR20080024530A (ko) 2008-03-18
DE112006001822T5 (de) 2008-05-21
CN101223521A (zh) 2008-07-16
JPWO2007010836A1 (ja) 2009-01-29
US20100076745A1 (en) 2010-03-25
CN101223521B (zh) 2010-06-16

Similar Documents

Publication Publication Date Title
KR101136007B1 (ko) 문서 감성 분석 시스템 및 그 방법
KR101339103B1 (ko) 의미적 자질을 이용한 문서 분류 시스템 및 그 방법
JP3429184B2 (ja) テキスト構造解析装置および抄録装置、並びにプログラム記録媒体
JP4634736B2 (ja) 専門的記述と非専門的記述間の語彙変換方法・プログラム・システム
CN103678316B (zh) 实体关系分类装置和实体关系分类方法
US20150100307A1 (en) Text segmentation with multiple granularity levels
CN104281645A (zh) 一种基于词汇语义和句法依存的情感关键句识别方法
Suba et al. Hybrid inflectional stemmer and rule-based derivational stemmer for gujarati
CN109298796B (zh) 一种词联想方法及装置
CN106446018B (zh) 基于人工智能的查询信息处理方法和装置
CN106570112A (zh) 基于改进的蚁群算法实现文本聚类
JP5718405B2 (ja) 発話選択装置、方法、及びプログラム、対話装置及び方法
CN113688624A (zh) 一种基于语言风格的人格预测方法及装置
Albared et al. Arabic term extraction using combined approach on Islamic document
WO2007010836A1 (ja) コミュニティ特有表現検出装置及び方法
JP2000259653A (ja) 音声認識装置及び音声認識方法
JP2005202924A (ja) 対訳判断装置、方法及びプログラム
CN113486155B (zh) 一种融合固定短语信息的中文命名方法
CN116362224A (zh) 用于多媒体作战救援的文本特征提取方法
Cholakov et al. Automated verb sense labelling based on linked lexical resources
Ablimit et al. Multilingual stemming and term extraction for Uyghur, Kazak and Kirghiz
Arumugam et al. Similitude Based Segment Graph Construction and Segment Ranking for Automatic Summarization of Text Document
JPH07325837A (ja) 抽象単語による通信文検索装置及び抽象単語による通信文検索方法
JP5860861B2 (ja) 焦点推定装置、モデル学習装置、方法、及びプログラム
Girault Concept lattice mining for unsupervised named entity annotation

Legal Events

Date Code Title Description
WWE Wipo information: entry into national phase

Ref document number: 200680025802.1

Country of ref document: CN

121 Ep: the epo has been informed by wipo that ep was designated in this application
WWE Wipo information: entry into national phase

Ref document number: 1120060018221

Country of ref document: DE

Ref document number: 2007525983

Country of ref document: JP

WWE Wipo information: entry into national phase

Ref document number: 1020087001074

Country of ref document: KR

WWE Wipo information: entry into national phase

Ref document number: 11990495

Country of ref document: US

RET De translation (de og part 6b)

Ref document number: 112006001822

Country of ref document: DE

Date of ref document: 20080521

Kind code of ref document: P

122 Ep: pct application non-entry in european phase

Ref document number: 06781076

Country of ref document: EP

Kind code of ref document: A1