WO2007010836A1 - コミュニティ特有表現検出装置及び方法 - Google Patents
コミュニティ特有表現検出装置及び方法 Download PDFInfo
- Publication number
- WO2007010836A1 WO2007010836A1 PCT/JP2006/314000 JP2006314000W WO2007010836A1 WO 2007010836 A1 WO2007010836 A1 WO 2007010836A1 JP 2006314000 W JP2006314000 W JP 2006314000W WO 2007010836 A1 WO2007010836 A1 WO 2007010836A1
- Authority
- WO
- WIPO (PCT)
- Prior art keywords
- word
- community
- gram
- significance
- selecting
- Prior art date
- Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
- Ceased
Links
Classifications
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/237—Lexical tools
- G06F40/242—Dictionaries
-
- G—PHYSICS
- G06—COMPUTING OR CALCULATING; COUNTING
- G06F—ELECTRIC DIGITAL DATA PROCESSING
- G06F40/00—Handling natural language data
- G06F40/20—Natural language analysis
- G06F40/279—Recognition of textual entities
- G06F40/289—Phrasal analysis, e.g. finite state techniques or chunking
Definitions
- the present invention relates to an apparatus and method for detecting a community-specific expression from expressions used in a community based on word formation theory.
- Patent Document 1 JP 2002-297589 “Unknown word collection method”
- Patent Document 2 JP-A-5-113997 “Dictionary Data Collection Device”
- Patent Document 3 JP 2004-265440 “Unknown Word Registration Device and Method and Storage Medium”
- Patent Document 4 JP 2005-309853 “Vocabulary Conversion Method Between Professional Description and Non-Professional Description 'Program' System”
- Non-patent document 1 Yuji Nakagawa, Yasuaki Yumoto, & Nada Nada (2003). Extraction of specialized terms based on appearance frequency and connection frequency. Natural language processing, 10 (1), 27-45.
- Non-Patent Literature 2 Zhaoqing University, & Fuyue Fumane (2004). Basic Research for Identifying New Words Important in Specialized Fields. Proc. Of the 10th Annual Conference of the Language Processing Society, (pp. 189 -191).
- Non-Patent Document 3 Satoshi Fujii, Katsunobu Ito, Tomoaki Akiba (2003), IPA Unexplored Software Creation Project “CYCLONE: Building the Strongest Dictionary Site”, www.ipa.go.jp/about/news/event/ pdf / 29A7_f ujii.pdf
- Non-patent document 4 Akihiko Yonekawa (1998) “Science of youth language” Tokyo: Meiji Shoin
- Patent Document 1 Japanese Patent Application Laid-Open No. 2002-297589 “Unknown Word Collection Method”
- Patent Document 3 Japanese Patent Application Laid-Open No. 2004-265440 “Unknown Word Registration Device”
- Patent Document 1 Japanese Patent Application Laid-Open No. 2002-297589 “Unknown Word Collection Method”
- Patent Document 1 This method also has the same power. Basically, many things that are not registered in the dictionary are collected by human stakes. In the detection of these unknown words, the target is almost limited to nouns, and rarely focus on the problem of collecting truly new expressions.
- Non-patent Document 4 In sociolinguistics, there is a field that collects and analyzes “young people” used by high school students and university students (Non-patent Document 4). Existing research on community-specific expressions seems to be close to the present invention, but in the field of sociolinguistics, it has been proposed that a method should be proposed for regularly collecting youth and buzzwords.
- Document gathering power used in a given community with the following means (a) to (d) A device that searches for expressions unique to a given community,
- the apparatus according to (1) further comprising means for collecting the document set by performing a data search using a term included in a predetermined term list as a keyword.
- the means for extracting the n-gram collocation uses a document used in a plurality of communities, and calculates the significance of the n-gram collocation used in the predetermined community and the n-gram collocation used in other communities.
- a method for retrieving an expression specific to a given community from a set of documents used in the given community comprising the following steps (a) to (d):
- the program according to (6) further comprising means for collecting the document set by searching data using a term included in a predetermined term list as a keyword.
- the invention of the present application is an extension of the language between main parts of speech and can be applied to other languages.
- the expression “He 747'ed to Chicago.” Is possible. This is a verbal version of the aircraft model. Also, "The web-logging is becoming a social phenomenon.” This is an example of a noun verb.
- FIG. 1 shows an example of a system when the present invention is implemented.
- a user PC 110 Connected to the network 140 are a user PC 110, a site server (1) 120, a site server (2) 130, and the like.
- the site server (1) 120, site server (2) 130, etc. connected to the network 140 are accessed, and necessary information is acquired using a search tool or the like.
- the present invention shows a search on the Internet as an embodiment, the present invention is not limited to this, and any other method can be applied as long as the system can search information.
- the acquired information can be processed by a computer program on the user PC to obtain the desired result.
- FIG. 2 shows a user PC that implements part of the present invention.
- the housing 200 includes a storage device 210, a main memory 220, an output device 230, a central control device (CPU) 240, an operation device 250, and a network 1/0260.
- the user operates the operation device 250 and obtains necessary information from each site on the Internet through the network I / O.
- the central controller 240 downloads the document processing program stored in the storage device 210 to the memory, performs predetermined data processing using information retrieved from the Internet, and displays the result on the output device 230. .
- FIG. 3 shows a block diagram of a community specific expression detection apparatus according to the present invention.
- 3 10 is a community document search unit
- 314 is a website
- 316 is a term list storage unit
- 320 is a document processing unit
- 330 is an n-gram collocation extraction unit
- 335 is a significance determination unit
- 340 is a word base selection unit
- 350 is The left and right extension part of the word base
- 354 is the left extension rule storage part
- 356 is the right extension rule storage part
- 360 is the new expression selection part
- 365 is the language rule storage part
- 370 is the output part. Details of these will be described below.
- Step 410 Collect documents for community use
- Step 420 n-gram collocation extraction
- Step 430 Selecting the core element (word base) of the new expression
- Step 440 Select extended word base
- Step 450 New expression selection
- Step 510 Get candidate documents by specifying terms
- Step 520 Preprocessing candidate documents
- Step 530 Remove noise document
- Step 540 Need to search for other community documents
- Step 510 Acquisition of candidate documents
- a term list including a predetermined term is used to collect documents used by parties in a predetermined community.
- the term list is stored in the term list storage unit (Fig. 3: 316).
- the term list is a set of terms that become keywords in one community. For example, if “wine lovers” is selected as one community, the component of the term list is “wine brands”. According to the brands listed in the wine terminology, use the Internet search tool to collect information about the wine ( Figure 3: 314). Here, brands such as “Hauslese”, “Chateau Kyule Bonn”, “Chateau Margoichi”, “Vine Santo Toscano” and the like can be designated. Candidate documents are searched from the database using this term as a keyword. Any database can be used as long as such information is stored in the database, but in this embodiment, a method for searching candidate documents using an Internet search engine will be described. [0017] (1 2) Step 520: Preprocessing of candidate document
- the web page information-powered document is first extracted and analyzed.
- segmentation is performed to extract content words, particles, auxiliary verbs, etc., and feature values representing the characteristics of these documents are obtained.
- feature values representing the characteristics of these documents are obtained.
- noise documents are removed as follows.
- Documents that automatically collect this information from Internet web pages contain a variety of information and are often not available as they are.
- documents corresponding to garbage documents, list documents, and diary documents are removed from these documents as noise documents.
- a document that satisfies all the conditions such as a document with a small number of content words or a document with a low proper noun ratio.
- the number of content words is the number of content words contained in a document described on one web page.
- Content words are words that correspond to nouns, verbs, adjectives, and adverbs, excluding particles and auxiliary verbs.
- the proper nouns mentioned here are nouns that are generally recognized as proper nouns.
- the proper noun ratio is the ratio between the number of proper nouns appearing on one web page and the number of content words.
- a document that satisfies all of the conditions such as a document having a high proper noun ratio, a document having a low correlation coefficient between the content word and the particle 'auxiliary verb', etc. is defined as a list information document. This is a document where information about objects in a certain area is stored as a simple list on an Internet site.
- a document to be defined is defined as a diary document. These are so-called These are documents that mainly contain other information such as documents used as personal diary writing sites and sites related to department stores. Based on the above definition, garbage documents, list documents, and diary documents are removed as noise documents.
- Step 540 Necessity of Search for Other Community Documents
- step 510 From step 510 to step 530, a set of documents used in a predetermined community is collected.
- step 540 a collection of documents used by other communities is collected as well.
- n-gram collocations word-level n-gram collocations (n-gram collocations) using statistical methods and those that appear significantly when used in a specific community. These are called community-specific collocations. These details will be described.
- An n-gram collocation is a sequence of one or more words: a unigram for one word, a bigram for two words, a trigram for three words. It is called (Tri-gram).
- Tri-gram bigrams and trigrams are used (FIG. 3: 330).
- the sample ratio is a ratio obtained from actual data
- pi and p2 are sample ratios.
- n-gram collocation W means to test whether it appears significantly biased towards the document in dl. Yes (one-sided test).
- null hypothesis the null hypothesis and the alternative hypothesis are as follows.
- a list of 2 grams and 3 grams appearing characteristically in a document set used by wine lovers and a document set used by sake lovers is extracted, and a Z test is performed.
- n-grams with a Z value of 1.65 or more are selected from a set of documents used by wine lovers.
- the n-gram extracted by the above method Take out the core element ( Figure 3: 340). To do this, break the n-gram chain for the time being and make a list of all the elements (morphemes) that occur there. From there, exclude those that are not likely to be core. Here, there is a function such as a particle, an auxiliary verb, a conjunction, a conjugation ending, and a break element such as “,”, “.”, “?”, Etc. as those that are not likely to be the core. Also excluded are “one hiragana character” and “one katakana character”. This creates a list of elements (the core list) that can be the core of the new expression.
- each word base candidate it is determined whether it is necessary to incorporate the preceding and succeeding elements based on the collocation pattern distribution (Fig. 3: 350).
- Z [X] is the Z value of the n-gram word group that we are currently focusing on.
- X be the core element
- [X + 1] be the element expanded by one word
- [X + 2] be the element expanded by two words.
- AvgZ ([X] [X + l]) is the word of all (n + 1) grams corresponding to [X] [X + 1] when expanded from the n-gram word base to the right It is the average of the base Z values (0 ⁇ Z;).
- Equation 6 is defined by taking the logarithm of Z.
- (ii) LZ> first threshold If it satisfies, it is selected as a candidate to expand to [X + 1] (610, 620, 650).
- the first threshold value is 5.0 in this embodiment, and Z ([X], [X + 1]) is represented by ([X], [X + 1]) (n + 1) Gram word base Z value of AvgZ ([X], [X + l], [X + 2]) is all (n + 2) grams corresponding to [X], [X + 1], [X + 2] This is the average of the Z values.
- the first threshold for LZ used in the first condition is set high. If this value is high, it will be judged that it can be recognized as a new expression enough even by judgment based on the value of Z. Therefore, it is selected as a possibility of new expression regardless of the value of Jratio (described later). To do.
- condition (ie) both conditions (i) and (ii) are met, it is selected as an expanded word candidate (650). If condition (i) is not met! /, It is not selected as a candidate for expansion (660). If the condition (i) is satisfied but the condition (ii) is not satisfied, the determination is made based on the second condition shown below (630, 640).
- the second threshold value for LZ used in the second condition is set to 3.0 in the example, and only when LZ is larger than this value and Jratio is 0.1 or more, new expression is possible. It is determined that there is sex.
- Nail is the number of (n + 2) grams corresponding to the target [X + 2].
- the elements of [X + 2], that is, “ga” and “ha” are called kOne elements. If there are multiple kOne elements as in this example, the average value of these Z values is calculated. In this case, since both are 2.00, the average value is 2.00.
- this kOne element is a “break element” indicating a break.
- a break element indicating a break.
- a grammatical break is shown.
- Jratio The proportion of kOne elements that are break elements is called Jratio.
- the left extension rule is explained using an example. Explain that [receiving] (Z value is 73.01) selected as a word base is extended to the left.
- Nounization and Examples include “base + suffix”, “verb conjunctive nounization”, “compound noun”, and the like. In each case, it is necessary to confirm the key to satisfy the rules for Japanese.
- the present invention can be applied not only to Japanese but also to foreign languages. I will explain using English as an example. Something that is used in English as a part of speech other than the original noun may be used as a noun. For example, it is made a noun by adding the following suffix. “Ness”: pleasantness, ugliness
- Verbification rules (step 720) Those that match the verbalization rules are also selected as candidates for word base expansion. Examples of verbs include “noun + do” and “general use of verb”. It is necessary to confirm whether the candidate selected for expansion satisfies the Japanese rules.
- a noun is combined with a verbal suffix such as “S”, “Buru”, or its conjugation, it is selected as a candidate for verbal expansion of the word base. For example, if “tea” is added to “tea” and “tea is made”, “beauty” is added to “beauty” by adding “bu”.
- An expanded word base is also selected as a candidate for expansion of the word base even if it is a general verb usage form excluding the form of “noun + verbal suffix”.
- verbs are added to the nouns and converted into verbs: “Demo, not demo, if demo”.
- new L ⁇ verbs such as “Gevaru, Hamoru, Tsumoru, Darguru” can be created in this way.
- the present invention can be applied not only to Japanese but also to foreign languages. I will explain using English as an example. Something that is originally used as a noun in English may be used as a verb. Are you googling?
- Step 710 to Step 740 If any of the above conditions from Step 710 to Step 740 is satisfied, it is selected as a candidate for expansion of the word base (760). If neither condition is met, it is not selected as a candidate for expansion of the word base (750).
- the LZ value is 3.01.
- expanded compound nouns include:
- FIG. 1 is a diagram showing an example of a system for carrying out the present invention.
- FIG. 2 is a block diagram of a PC that implements part of the present invention.
- FIG. 3 is a block diagram of a community specific expression detection device according to the present invention.
- FIG. 4 is a flowchart of the present invention.
- FIG. 5 is a flowchart of document collection according to the present invention.
- FIG. 6 is a flowchart for determining the suitability of an expanded word base.
- CPU Central control unit
Landscapes
- Engineering & Computer Science (AREA)
- Theoretical Computer Science (AREA)
- Health & Medical Sciences (AREA)
- Artificial Intelligence (AREA)
- Audiology, Speech & Language Pathology (AREA)
- Computational Linguistics (AREA)
- General Health & Medical Sciences (AREA)
- Physics & Mathematics (AREA)
- General Engineering & Computer Science (AREA)
- General Physics & Mathematics (AREA)
- Machine Translation (AREA)
- Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
Abstract
Description
Claims
Priority Applications (4)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| CN2006800258021A CN101223521B (zh) | 2005-07-15 | 2006-07-13 | 社群特有表现检测装置及方法 |
| JP2007525983A JPWO2007010836A1 (ja) | 2005-07-15 | 2006-07-13 | コミュニティ特有表現検出装置及び方法 |
| US11/990,495 US20100076745A1 (en) | 2005-07-15 | 2006-07-13 | Apparatus and Method of Detecting Community-Specific Expression |
| DE112006001822T DE112006001822T5 (de) | 2005-07-15 | 2006-07-13 | Vorrichtung und Verfahren zum Erfassen eines gemeinschaftsspezifischen Ausdrucks |
Applications Claiming Priority (2)
| Application Number | Priority Date | Filing Date | Title |
|---|---|---|---|
| JP2005207810 | 2005-07-15 | ||
| JP2005-207810 | 2005-07-15 |
Publications (1)
| Publication Number | Publication Date |
|---|---|
| WO2007010836A1 true WO2007010836A1 (ja) | 2007-01-25 |
Family
ID=37668717
Family Applications (1)
| Application Number | Title | Priority Date | Filing Date |
|---|---|---|---|
| PCT/JP2006/314000 Ceased WO2007010836A1 (ja) | 2005-07-15 | 2006-07-13 | コミュニティ特有表現検出装置及び方法 |
Country Status (6)
| Country | Link |
|---|---|
| US (1) | US20100076745A1 (ja) |
| JP (1) | JPWO2007010836A1 (ja) |
| KR (1) | KR20080024530A (ja) |
| CN (1) | CN101223521B (ja) |
| DE (1) | DE112006001822T5 (ja) |
| WO (1) | WO2007010836A1 (ja) |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010160534A (ja) * | 2009-01-06 | 2010-07-22 | Yahoo Japan Corp | 地域特性辞書生成方法及び装置 |
| JP7557770B2 (ja) | 2020-06-05 | 2024-09-30 | 国立大学法人北海道国立大学機構 | 専門用語抽出装置、専門用語抽出方法及びプログラム |
Families Citing this family (4)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US8473279B2 (en) * | 2008-05-30 | 2013-06-25 | Eiman Al-Shammari | Lemmatizing, stemming, and query expansion method and system |
| US8423350B1 (en) * | 2009-05-21 | 2013-04-16 | Google Inc. | Segmenting text for searching |
| US20110082687A1 (en) * | 2009-10-05 | 2011-04-07 | Marcelo Pham | Method and system for taking actions based on analysis of enterprise communication messages |
| KR101706827B1 (ko) * | 2014-12-04 | 2017-02-16 | 강원대학교산학협력단 | 개체 간 사회 관계 추출 장치 및 방법 |
Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH1185761A (ja) * | 1997-09-03 | 1999-03-30 | Ee I Soft Kk | 未知語登録装置および方法並びに記録媒体 |
| JP2004062262A (ja) * | 2002-07-25 | 2004-02-26 | Hitachi Ltd | 未知語を自動的に辞書へ登録する方法 |
Family Cites Families (22)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| US5265065A (en) * | 1991-10-08 | 1993-11-23 | West Publishing Company | Method and apparatus for information retrieval from a database by replacing domain specific stemmed phases in a natural language to create a search query |
| US5799268A (en) * | 1994-09-28 | 1998-08-25 | Apple Computer, Inc. | Method for extracting knowledge from online documentation and creating a glossary, index, help database or the like |
| US5704060A (en) * | 1995-05-22 | 1997-12-30 | Del Monte; Michael G. | Text storage and retrieval system and method |
| US6173298B1 (en) * | 1996-09-17 | 2001-01-09 | Asap, Ltd. | Method and apparatus for implementing a dynamic collocation dictionary |
| US5933822A (en) * | 1997-07-22 | 1999-08-03 | Microsoft Corporation | Apparatus and methods for an information retrieval system that employs natural language processing of search results to improve overall precision |
| GB2338089A (en) * | 1998-06-02 | 1999-12-08 | Sharp Kk | Indexing method |
| US6347316B1 (en) * | 1998-12-14 | 2002-02-12 | International Business Machines Corporation | National language proxy file save and incremental cache translation option for world wide web documents |
| US6442524B1 (en) * | 1999-01-29 | 2002-08-27 | Sony Corporation | Analyzing inflectional morphology in a spoken language translation system |
| US6356865B1 (en) * | 1999-01-29 | 2002-03-12 | Sony Corporation | Method and apparatus for performing spoken language translation |
| US7865358B2 (en) * | 2000-06-26 | 2011-01-04 | Oracle International Corporation | Multi-user functionality for converting data from a first form to a second form |
| US7225199B1 (en) * | 2000-06-26 | 2007-05-29 | Silver Creek Systems, Inc. | Normalizing and classifying locale-specific information |
| US8396859B2 (en) * | 2000-06-26 | 2013-03-12 | Oracle International Corporation | Subject matter context search engine |
| US6675159B1 (en) * | 2000-07-27 | 2004-01-06 | Science Applic Int Corp | Concept-based search and retrieval system |
| US7526425B2 (en) * | 2001-08-14 | 2009-04-28 | Evri Inc. | Method and system for extending keyword searching to syntactically and semantically annotated data |
| WO2005024604A2 (en) * | 2003-09-09 | 2005-03-17 | Siftology, Inc. | Dynamic lexicon |
| US20050149510A1 (en) * | 2004-01-07 | 2005-07-07 | Uri Shafrir | Concept mining and concept discovery-semantic search tool for large digital databases |
| US7260568B2 (en) * | 2004-04-15 | 2007-08-21 | Microsoft Corporation | Verifying relevance between keywords and web site contents |
| US20070217693A1 (en) * | 2004-07-02 | 2007-09-20 | Texttech, Llc | Automated evaluation systems & methods |
| US7571157B2 (en) * | 2004-12-29 | 2009-08-04 | Aol Llc | Filtering search results |
| WO2006096260A2 (en) * | 2005-01-31 | 2006-09-14 | Musgrove Technology Enterprises, Llc | System and method for generating an interlinked taxonomy structure |
| US7657421B2 (en) * | 2006-06-28 | 2010-02-02 | International Business Machines Corporation | System and method for identifying and defining idioms |
| US7698328B2 (en) * | 2006-08-11 | 2010-04-13 | Apple Inc. | User-directed search refinement |
-
2006
- 2006-07-13 WO PCT/JP2006/314000 patent/WO2007010836A1/ja not_active Ceased
- 2006-07-13 CN CN2006800258021A patent/CN101223521B/zh not_active Expired - Fee Related
- 2006-07-13 KR KR1020087001074A patent/KR20080024530A/ko not_active Ceased
- 2006-07-13 DE DE112006001822T patent/DE112006001822T5/de not_active Withdrawn
- 2006-07-13 US US11/990,495 patent/US20100076745A1/en not_active Abandoned
- 2006-07-13 JP JP2007525983A patent/JPWO2007010836A1/ja not_active Withdrawn
Patent Citations (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JPH1185761A (ja) * | 1997-09-03 | 1999-03-30 | Ee I Soft Kk | 未知語登録装置および方法並びに記録媒体 |
| JP2004062262A (ja) * | 2002-07-25 | 2004-02-26 | Hitachi Ltd | 未知語を自動的に辞書へ登録する方法 |
Non-Patent Citations (2)
| Title |
|---|
| MORI S. ET AL.: "n Glam Tokei ni yoru Corpus kara no Michigo Chushutsu", IEICE TECHNICAL REPORT NLC 95-8, vol. 95, no. 168, 20 July 1995 (1995-07-20), pages 7 - 12, XP003007716 * |
| NAGAO M. ET AL.: "Daikibo Nihongo Text no n glam Tokei no Tsukurikata to Goku no Jido Chushutsu 93-NL-96-1", vol. 93, no. 61, 9 July 1993 (1993-07-09), pages 1 - 8, XP003007717 * |
Cited By (2)
| Publication number | Priority date | Publication date | Assignee | Title |
|---|---|---|---|---|
| JP2010160534A (ja) * | 2009-01-06 | 2010-07-22 | Yahoo Japan Corp | 地域特性辞書生成方法及び装置 |
| JP7557770B2 (ja) | 2020-06-05 | 2024-09-30 | 国立大学法人北海道国立大学機構 | 専門用語抽出装置、専門用語抽出方法及びプログラム |
Also Published As
| Publication number | Publication date |
|---|---|
| KR20080024530A (ko) | 2008-03-18 |
| DE112006001822T5 (de) | 2008-05-21 |
| CN101223521A (zh) | 2008-07-16 |
| JPWO2007010836A1 (ja) | 2009-01-29 |
| US20100076745A1 (en) | 2010-03-25 |
| CN101223521B (zh) | 2010-06-16 |
Similar Documents
| Publication | Publication Date | Title |
|---|---|---|
| KR101136007B1 (ko) | 문서 감성 분석 시스템 및 그 방법 | |
| KR101339103B1 (ko) | 의미적 자질을 이용한 문서 분류 시스템 및 그 방법 | |
| JP3429184B2 (ja) | テキスト構造解析装置および抄録装置、並びにプログラム記録媒体 | |
| JP4634736B2 (ja) | 専門的記述と非専門的記述間の語彙変換方法・プログラム・システム | |
| CN103678316B (zh) | 实体关系分类装置和实体关系分类方法 | |
| US20150100307A1 (en) | Text segmentation with multiple granularity levels | |
| CN104281645A (zh) | 一种基于词汇语义和句法依存的情感关键句识别方法 | |
| Suba et al. | Hybrid inflectional stemmer and rule-based derivational stemmer for gujarati | |
| CN109298796B (zh) | 一种词联想方法及装置 | |
| CN106446018B (zh) | 基于人工智能的查询信息处理方法和装置 | |
| CN106570112A (zh) | 基于改进的蚁群算法实现文本聚类 | |
| JP5718405B2 (ja) | 発話選択装置、方法、及びプログラム、対話装置及び方法 | |
| CN113688624A (zh) | 一种基于语言风格的人格预测方法及装置 | |
| Albared et al. | Arabic term extraction using combined approach on Islamic document | |
| WO2007010836A1 (ja) | コミュニティ特有表現検出装置及び方法 | |
| JP2000259653A (ja) | 音声認識装置及び音声認識方法 | |
| JP2005202924A (ja) | 対訳判断装置、方法及びプログラム | |
| CN113486155B (zh) | 一种融合固定短语信息的中文命名方法 | |
| CN116362224A (zh) | 用于多媒体作战救援的文本特征提取方法 | |
| Cholakov et al. | Automated verb sense labelling based on linked lexical resources | |
| Ablimit et al. | Multilingual stemming and term extraction for Uyghur, Kazak and Kirghiz | |
| Arumugam et al. | Similitude Based Segment Graph Construction and Segment Ranking for Automatic Summarization of Text Document | |
| JPH07325837A (ja) | 抽象単語による通信文検索装置及び抽象単語による通信文検索方法 | |
| JP5860861B2 (ja) | 焦点推定装置、モデル学習装置、方法、及びプログラム | |
| Girault | Concept lattice mining for unsupervised named entity annotation |
Legal Events
| Date | Code | Title | Description |
|---|---|---|---|
| WWE | Wipo information: entry into national phase |
Ref document number: 200680025802.1 Country of ref document: CN |
|
| 121 | Ep: the epo has been informed by wipo that ep was designated in this application | ||
| WWE | Wipo information: entry into national phase |
Ref document number: 1120060018221 Country of ref document: DE Ref document number: 2007525983 Country of ref document: JP |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 1020087001074 Country of ref document: KR |
|
| WWE | Wipo information: entry into national phase |
Ref document number: 11990495 Country of ref document: US |
|
| RET | De translation (de og part 6b) |
Ref document number: 112006001822 Country of ref document: DE Date of ref document: 20080521 Kind code of ref document: P |
|
| 122 | Ep: pct application non-entry in european phase |
Ref document number: 06781076 Country of ref document: EP Kind code of ref document: A1 |