EA200000321A1 - Система автоматической идентификации языка для многоязычного оптического распознавания символов - Google Patents

Система автоматической идентификации языка для многоязычного оптического распознавания символов

Info

Publication number
EA200000321A1
EA200000321A1 EA200000321A EA200000321A EA200000321A1 EA 200000321 A1 EA200000321 A1 EA 200000321A1 EA 200000321 A EA200000321 A EA 200000321A EA 200000321 A EA200000321 A EA 200000321A EA 200000321 A1 EA200000321 A1 EA 200000321A1
Authority
EA
Eurasian Patent Office
Prior art keywords
language
zone
identification system
automatic identification
regions
Prior art date
Application number
EA200000321A
Other languages
English (en)
Other versions
EA001689B1 (ru
Inventor
Леонард К. Пон
Тапас Канунго
Дзун Янг
Кеннет Чан Чой
Минди Р. Боксер
Original Assignee
Каер Корпорейшн
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Каер Корпорейшн filed Critical Каер Корпорейшн
Publication of EA200000321A1 publication Critical patent/EA200000321A1/ru
Publication of EA001689B1 publication Critical patent/EA001689B1/ru

Links

Classifications

    • GPHYSICS
    • G06COMPUTING; CALCULATING OR COUNTING
    • G06VIMAGE OR VIDEO RECOGNITION OR UNDERSTANDING
    • G06V30/00Character recognition; Recognising digital ink; Document-oriented image-based pattern recognition
    • G06V30/10Character recognition
    • G06V30/24Character recognition characterised by the processing or recognition method
    • G06V30/242Division of the character sequences into groups prior to recognition; Selection of dictionaries

Landscapes

  • Engineering & Computer Science (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Physics & Mathematics (AREA)
  • General Physics & Mathematics (AREA)
  • Multimedia (AREA)
  • Theoretical Computer Science (AREA)
  • Character Discrimination (AREA)
  • Machine Translation (AREA)

Abstract

В данном изобретении применяют словарный подход для идентификации языков в различных зонах многоязычного документа. На первом этапе образ документа сегментируют на различные зоны, области и словоформы, с использованием подходящих геометрических свойств. В каждой зоне словоформы сравнивают со словарями, сопоставляемыми различным языкам-кандидатам, и язык, который проявляет наивысший показатель доверительности, первоначально идентифицируют в качестве языка данной зоны. Затем каждую зону расщепляют на области. После этого производят идентификацию языка каждой области с использованием показателей доверительности для слов данной области. Для любого определения языка, имеющего низкое значение доверительности, ранее определенный язык зоны применяют с целью способствовать процессу идентификации.Международная заявка была опубликована вместе с отчетом о международном поиске.
EA200000321A 1997-09-15 1997-11-20 Система автоматической идентификации языка для многоязычного оптического распознавания символов EA001689B1 (ru)

Applications Claiming Priority (2)

Application Number Priority Date Filing Date Title
US08/929,788 US6047251A (en) 1997-09-15 1997-09-15 Automatic language identification system for multilingual optical character recognition
PCT/US1997/018705 WO1999014708A1 (en) 1997-09-15 1997-11-20 Automatic language identification system for multilingual optical character recognition

Publications (2)

Publication Number Publication Date
EA200000321A1 true EA200000321A1 (ru) 2000-10-30
EA001689B1 EA001689B1 (ru) 2001-06-25

Family

ID=25458457

Family Applications (1)

Application Number Title Priority Date Filing Date
EA200000321A EA001689B1 (ru) 1997-09-15 1997-11-20 Система автоматической идентификации языка для многоязычного оптического распознавания символов

Country Status (8)

Country Link
US (1) US6047251A (ru)
EP (1) EP1016033B1 (ru)
CN (1) CN1122243C (ru)
AT (1) ATE243342T1 (ru)
AU (1) AU5424498A (ru)
DE (1) DE69722971T2 (ru)
EA (1) EA001689B1 (ru)
WO (1) WO1999014708A1 (ru)

Families Citing this family (80)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US6449718B1 (en) * 1999-04-09 2002-09-10 Xerox Corporation Methods and apparatus for partial encryption of tokenized documents
US20020023123A1 (en) * 1999-07-26 2002-02-21 Justin P. Madison Geographic data locator
EP2448155A3 (en) 1999-11-10 2014-05-07 Pandora Media, Inc. Internet radio and broadcast method
US6389467B1 (en) 2000-01-24 2002-05-14 Friskit, Inc. Streaming media search and continuous playback system of media resources located by multiple network addresses
US6567801B1 (en) 2000-03-16 2003-05-20 International Business Machines Corporation Automatically initiating a knowledge portal query from within a displayed document
US6584469B1 (en) * 2000-03-16 2003-06-24 International Business Machines Corporation Automatically initiating a knowledge portal query from within a displayed document
EP1139231A1 (en) * 2000-03-31 2001-10-04 Fujitsu Limited Document processing apparatus and method
US6738745B1 (en) * 2000-04-07 2004-05-18 International Business Machines Corporation Methods and apparatus for identifying a non-target language in a speech recognition system
US7024485B2 (en) * 2000-05-03 2006-04-04 Yahoo! Inc. System for controlling and enforcing playback restrictions for a media file by splitting the media file into usable and unusable portions for playback
US7251665B1 (en) 2000-05-03 2007-07-31 Yahoo! Inc. Determining a known character string equivalent to a query string
US8352331B2 (en) 2000-05-03 2013-01-08 Yahoo! Inc. Relationship discovery engine
US7162482B1 (en) * 2000-05-03 2007-01-09 Musicmatch, Inc. Information retrieval engine
US6678415B1 (en) * 2000-05-12 2004-01-13 Xerox Corporation Document image decoding using an integrated stochastic language model
DE10196421T5 (de) * 2000-07-11 2006-07-13 Launch Media, Inc., Santa Monica Online Playback-System mit Gemeinschatsausrichtung
US8271333B1 (en) 2000-11-02 2012-09-18 Yahoo! Inc. Content-related wallpaper
US7493250B2 (en) * 2000-12-18 2009-02-17 Xerox Corporation System and method for distributing multilingual documents
US7406529B2 (en) * 2001-02-09 2008-07-29 Yahoo! Inc. System and method for detecting and verifying digitized content over a computer network
US7574513B2 (en) 2001-04-30 2009-08-11 Yahoo! Inc. Controllable track-skipping
GB0111012D0 (en) 2001-05-04 2001-06-27 Nokia Corp A communication terminal having a predictive text editor application
DE10126835B4 (de) * 2001-06-01 2004-04-29 Siemens Dematic Ag Verfahren und Vorrichtung zum automatischen Lesen von Adressen in mehr als einer Sprache
US7191116B2 (en) * 2001-06-19 2007-03-13 Oracle International Corporation Methods and systems for determining a language of a document
US7707221B1 (en) 2002-04-03 2010-04-27 Yahoo! Inc. Associating and linking compact disc metadata
US7020338B1 (en) * 2002-04-08 2006-03-28 The United States Of America As Represented By The National Security Agency Method of identifying script of line of text
US7305483B2 (en) 2002-04-25 2007-12-04 Yahoo! Inc. Method for the real-time distribution of streaming data on a network
RU2251737C2 (ru) * 2002-10-18 2005-05-10 Аби Софтвер Лтд. Способ автоматического определения языка распознаваемого текста при многоязычном распознавании
JP3919617B2 (ja) * 2002-07-09 2007-05-30 キヤノン株式会社 文字認識装置および文字認識方法、プログラムおよび記憶媒体
US6669085B1 (en) * 2002-08-07 2003-12-30 Hewlett-Packard Development Company, L.P. Making language localization and telecommunications settings in a multi-function device through image scanning
US20040078191A1 (en) * 2002-10-22 2004-04-22 Nokia Corporation Scalable neural network-based language identification from written text
FR2848688A1 (fr) * 2002-12-17 2004-06-18 France Telecom Identification de langue d'un texte
US7672873B2 (en) * 2003-09-10 2010-03-02 Yahoo! Inc. Music purchasing and playing system and method
US7424672B2 (en) * 2003-10-03 2008-09-09 Hewlett-Packard Development Company, L.P. System and method of specifying image document layout definition
JP3890326B2 (ja) * 2003-11-07 2007-03-07 キヤノン株式会社 情報処理装置、情報処理方法ならびに記録媒体、プログラム
US8027832B2 (en) * 2005-02-11 2011-09-27 Microsoft Corporation Efficient language identification
JP4311365B2 (ja) * 2005-03-25 2009-08-12 富士ゼロックス株式会社 文書処理装置およびプログラム
JP4856925B2 (ja) * 2005-10-07 2012-01-18 株式会社リコー 画像処理装置、画像処理方法及び画像処理プログラム
US8185376B2 (en) * 2006-03-20 2012-05-22 Microsoft Corporation Identifying language origin of words
US7493293B2 (en) * 2006-05-31 2009-02-17 International Business Machines Corporation System and method for extracting entities of interest from text using n-gram models
US8140267B2 (en) * 2006-06-30 2012-03-20 International Business Machines Corporation System and method for identifying similar molecules
US9020811B2 (en) * 2006-10-13 2015-04-28 Syscom, Inc. Method and system for converting text files searchable text and for processing the searchable text
US7912289B2 (en) 2007-05-01 2011-03-22 Microsoft Corporation Image text replacement
US9141607B1 (en) * 2007-05-30 2015-09-22 Google Inc. Determining optical character recognition parameters
GB0717067D0 (en) * 2007-09-03 2007-10-10 Ibm An Apparatus for preparing a display document for analysis
US8233726B1 (en) * 2007-11-27 2012-07-31 Googe Inc. Image-domain script and language identification
US8107671B2 (en) * 2008-06-26 2012-01-31 Microsoft Corporation Script detection service
US8073680B2 (en) * 2008-06-26 2011-12-06 Microsoft Corporation Language detection service
US8266514B2 (en) * 2008-06-26 2012-09-11 Microsoft Corporation Map service
US8019596B2 (en) * 2008-06-26 2011-09-13 Microsoft Corporation Linguistic service platform
US8224641B2 (en) * 2008-11-19 2012-07-17 Stratify, Inc. Language identification for documents containing multiple languages
US8224642B2 (en) * 2008-11-20 2012-07-17 Stratify, Inc. Automated identification of documents as not belonging to any language
CN101751567B (zh) * 2008-12-12 2012-10-17 汉王科技股份有限公司 快速文本识别方法
US8326602B2 (en) * 2009-06-05 2012-12-04 Google Inc. Detecting writing systems and languages
US8468011B1 (en) * 2009-06-05 2013-06-18 Google Inc. Detecting writing systems and languages
CN102024138B (zh) * 2009-09-15 2013-01-23 富士通株式会社 字符识别方法和字符识别装置
US8756215B2 (en) * 2009-12-02 2014-06-17 International Business Machines Corporation Indexing documents
US20120035905A1 (en) * 2010-08-09 2012-02-09 Xerox Corporation System and method for handling multiple languages in text
US8635061B2 (en) * 2010-10-14 2014-01-21 Microsoft Corporation Language identification in multilingual text
JP5672003B2 (ja) * 2010-12-28 2015-02-18 富士通株式会社 文字認識処理装置及びプログラム
US8600730B2 (en) * 2011-02-08 2013-12-03 Microsoft Corporation Language segmentation of multilingual texts
CN102156889A (zh) * 2011-03-31 2011-08-17 汉王科技股份有限公司 一种识别手写文本行语言类别的方法及装置
US9519641B2 (en) * 2012-09-18 2016-12-13 Abbyy Development Llc Photography recognition translation
KR101686363B1 (ko) 2012-10-10 2016-12-13 모토로라 솔루션즈, 인크. 문서에 사용된 언어를 식별하고, 식별된 언어에 기초하여 ocr 인식을 수행하는 방법 및 장치
US9411801B2 (en) * 2012-12-21 2016-08-09 Abbyy Development Llc General dictionary for all languages
CN103902993A (zh) * 2012-12-28 2014-07-02 佳能株式会社 文档图像识别方法和设备
US9269352B2 (en) * 2013-05-13 2016-02-23 GM Global Technology Operations LLC Speech recognition with a plurality of microphones
CN103285360B (zh) * 2013-06-09 2014-09-17 王京涛 一种治疗血栓闭塞性脉管炎的中药制剂及其制备方法
BR112016002229A2 (pt) 2013-08-09 2017-08-01 Behavioral Recognition Sys Inc sistema de reconhecimento de comportamento neurolinguístico cognitivo para fusão de dados de multissensor
RU2613847C2 (ru) 2013-12-20 2017-03-21 ООО "Аби Девелопмент" Выявление китайской, японской и корейской письменности
JP2015210683A (ja) * 2014-04-25 2015-11-24 株式会社リコー 情報処理システム、情報処理装置、情報処理方法およびプログラム
US9798943B2 (en) * 2014-06-09 2017-10-24 I.R.I.S. Optical character recognition method
WO2016017009A1 (ja) * 2014-07-31 2016-02-04 楽天株式会社 メッセージ処理装置、メッセージ処理方法、記録媒体およびプログラム
US10963651B2 (en) * 2015-06-05 2021-03-30 International Business Machines Corporation Reformatting of context sensitive data
JP6655331B2 (ja) * 2015-09-24 2020-02-26 Dynabook株式会社 電子機器及び方法
CN106598937B (zh) * 2015-10-16 2019-10-18 阿里巴巴集团控股有限公司 用于文本的语种识别方法、装置和电子设备
CN107092903A (zh) * 2016-02-18 2017-08-25 阿里巴巴集团控股有限公司 信息识别方法及装置
US10311330B2 (en) 2016-08-17 2019-06-04 International Business Machines Corporation Proactive input selection for improved image analysis and/or processing workflows
US10579741B2 (en) 2016-08-17 2020-03-03 International Business Machines Corporation Proactive input selection for improved machine translation
US10460192B2 (en) * 2016-10-21 2019-10-29 Xerox Corporation Method and system for optical character recognition (OCR) of multi-language content
US10579733B2 (en) 2018-05-10 2020-03-03 Google Llc Identifying codemixed text
US11720752B2 (en) * 2020-07-07 2023-08-08 Sap Se Machine learning enabled text analysis with multi-language support
WO2021081562A2 (en) * 2021-01-20 2021-04-29 Innopeak Technology, Inc. Multi-head text recognition model for multi-lingual optical character recognition

Family Cites Families (7)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US3988715A (en) * 1975-10-24 1976-10-26 International Business Machines Corporation Multi-channel recognition discriminator
US4829580A (en) * 1986-03-26 1989-05-09 Telephone And Telegraph Company, At&T Bell Laboratories Text analysis system with letter sequence recognition and speech stress assignment arrangement
US5062143A (en) * 1990-02-23 1991-10-29 Harris Corporation Trigram-based method of language identification
US5182708A (en) * 1990-12-11 1993-01-26 Ricoh Corporation Method and apparatus for classifying text
US5371807A (en) * 1992-03-20 1994-12-06 Digital Equipment Corporation Method and apparatus for text classification
GB9220404D0 (en) * 1992-08-20 1992-11-11 Nat Security Agency Method of identifying,retrieving and sorting documents
US5548507A (en) * 1994-03-14 1996-08-20 International Business Machines Corporation Language identification process using coded language words

Also Published As

Publication number Publication date
CN1276077A (zh) 2000-12-06
EP1016033B1 (en) 2003-06-18
EP1016033A1 (en) 2000-07-05
WO1999014708A1 (en) 1999-03-25
US6047251A (en) 2000-04-04
ATE243342T1 (de) 2003-07-15
DE69722971T2 (de) 2003-12-04
CN1122243C (zh) 2003-09-24
AU5424498A (en) 1999-04-05
DE69722971D1 (de) 2003-07-24
EA001689B1 (ru) 2001-06-25

Similar Documents

Publication Publication Date Title
EA200000321A1 (ru) Система автоматической идентификации языка для многоязычного оптического распознавания символов
US20150186361A1 (en) Method and apparatus for improving a bilingual corpus, machine translation method and apparatus
WO2000033211A3 (en) Automatic segmentation of a text
Hellinger English–Gender in a global language
BR9914551A (pt) Processo e sistema para macro-linguagem extensìvel
CN102023972A (zh) 基于结构化的翻译记忆的自动翻译系统及其自动翻译方法
EA200301188A1 (ru) Способ и средства преобразования контента
Kosyreva et al. Axioms of interlinguistics in the context of language globalization
CN109543023B (zh) 基于trie和LCS算法的文献分类方法和系统
Tran et al. Word re-segmentation in Chinese-Vietnamese machine translation
Seddah et al. Ubiquitous usage of a french large corpus: Processing the est republicain corpus
Lee et al. papago: A machine translation service with word sense disambiguation and currency conversion
Kawtrakul et al. Backward transliteration for Thai document retrieval
Meelen et al. Towards a historical treebank of Middle and Early Modern Welsh, part I: Workflow and POS tagging
Akhtamova Phraseological expressions in the modern English language and their derivative features
Doermann et al. Translation lexicon acquisition from bilingual dictionaries
Mortensen Hmong-Mien languages
Beddows Translations and adaptations in Francophone Canada
Boizou et al. An online linguistic analyser for scottish gaelic
Osenova et al. Learning a token classification from a large corpus.(A case study in abbreviations)
KR100204068B1 (ko) 개념기반 다국어 번역시스템의 문법 자동수정 방법
Lehal et al. A Hindi to Urdu transliteration system
Goodridge The Dangers of Exclusionary Humanism: Advocating the Functionalist Paradigm Alongside Literal Translation in the Translation of Postcolonial Literature: Angela Carter’s ‘Our Lady of the Massacre’
Rustamova THE ANALYTICAL WORD FORMATION IN COMPARISON FRENCH AND RUSSIAN LANGUAGES
Wills et al. Integrating TEI/XML Text with Semantic Lexicographic Data.

Legal Events

Date Code Title Description
MM4A Lapse of a eurasian patent due to non-payment of renewal fees within the time limit in the following designated state(s)

Designated state(s): AM AZ BY KZ KG MD TJ TM RU