JP2024539003A - 事前トレーニングされた言語モデルの単一のトランスフォーマ層からのマルチヘッドネットワークの微調整 - Google Patents

事前トレーニングされた言語モデルの単一のトランスフォーマ層からのマルチヘッドネットワークの微調整 Download PDF

Info

Publication number
JP2024539003A
JP2024539003A JP2024522110A JP2024522110A JP2024539003A JP 2024539003 A JP2024539003 A JP 2024539003A JP 2024522110 A JP2024522110 A JP 2024522110A JP 2024522110 A JP2024522110 A JP 2024522110A JP 2024539003 A JP2024539003 A JP 2024539003A
Authority
JP
Japan
Prior art keywords
layers
machine learning
learning model
fine
utterance
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
JP2024522110A
Other languages
English (en)
Japanese (ja)
Other versions
JPWO2023064033A5 (https=
JP2024539003A5 (https=
Inventor
ブー,タン・ティエン
ファム,トゥエン・クアン
ネザミ,オミッド・モハマド
ジョンソン,マーク・エドワード
ドゥオング,タン・ロング
ホアン,コン・ズイ・ブー
Original Assignee
オラクル・インターナショナル・コーポレイション
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by オラクル・インターナショナル・コーポレイション filed Critical オラクル・インターナショナル・コーポレイション
Publication of JP2024539003A publication Critical patent/JP2024539003A/ja
Publication of JPWO2023064033A5 publication Critical patent/JPWO2023064033A5/ja
Publication of JP2024539003A5 publication Critical patent/JP2024539003A5/ja
Pending legal-status Critical Current

Links

Images

Classifications

    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/096Transfer learning
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L51/00User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
    • H04L51/02User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail using automatic reactions or user delegation, e.g. automatic replies or chatbot-generated messages
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • G10L2015/0635Training updating or merging of old and new templates; Mean values; Weighting
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/223Execution procedure of a spoken command

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Artificial Intelligence (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Medical Informatics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
JP2024522110A 2021-10-12 2022-08-17 事前トレーニングされた言語モデルの単一のトランスフォーマ層からのマルチヘッドネットワークの微調整 Pending JP2024539003A (ja)

Applications Claiming Priority (5)

Application Number Priority Date Filing Date Title
US202163254740P 2021-10-12 2021-10-12
US63/254,740 2021-10-12
US17/735,651 2022-05-03
US17/735,651 US12512091B2 (en) 2021-10-12 2022-05-03 Fine-tuning multi-head network from a single transformer layer of pre-trained language model
PCT/US2022/040530 WO2023064033A1 (en) 2021-10-12 2022-08-17 Fine-tuning multi-head network from a single transformer layer of pre-trained language model

Publications (3)

Publication Number Publication Date
JP2024539003A true JP2024539003A (ja) 2024-10-28
JPWO2023064033A5 JPWO2023064033A5 (https=) 2025-08-04
JP2024539003A5 JP2024539003A5 (https=) 2025-08-04

Family

ID=85798249

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2024522110A Pending JP2024539003A (ja) 2021-10-12 2022-08-17 事前トレーニングされた言語モデルの単一のトランスフォーマ層からのマルチヘッドネットワークの微調整

Country Status (6)

Country Link
US (2) US12512091B2 (https=)
JP (1) JP2024539003A (https=)
KR (1) KR20240089615A (https=)
CN (1) CN118140230A (https=)
GB (1) GB2631139A (https=)
WO (1) WO2023064033A1 (https=)

Families Citing this family (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12548552B2 (en) * 2021-11-19 2026-02-10 International Business Machines Corporation Dynamic language selection of an AI voice assistance system
US11947935B2 (en) * 2021-11-24 2024-04-02 Microsoft Technology Licensing, Llc. Custom models for source code generation via prefix-tuning
US20240061835A1 (en) * 2022-08-22 2024-02-22 Oracle International Corporation System and method of selective fine-tuning for custom training of a natural language to logical form model
US20240169165A1 (en) * 2022-11-17 2024-05-23 Samsung Electronics Co., Ltd. Automatically Generating Annotated Ground-Truth Corpus for Training NLU Model
US12562163B2 (en) * 2023-05-12 2026-02-24 Servicenow, Inc. Bidirectional assistant for development platforms
CN116774140A (zh) * 2023-06-26 2023-09-19 南京邮电大学 基于残差注意力网络的无网格信号源doa估计方法
US20250005282A1 (en) * 2023-06-29 2025-01-02 Amazon Technologies, Inc. Domain entity extraction for performing text analysis tasks
CN118446218B (zh) * 2024-05-16 2024-11-01 西南交通大学 一种对抗式阅读理解嵌套命名实体识别方法
CA3253531A1 (en) * 2024-06-14 2026-01-19 The Toronto-Dominion Bank Context retrieval for in-context learning model
WO2026000314A1 (en) * 2024-06-27 2026-01-02 Beijing Youzhuju Network Technology Co., Ltd. Model-based task processing
JP7658644B1 (ja) * 2024-10-21 2025-04-08 スパーブエーアイ カンパニー リミテッド 事前学習されたベースモデルに基づいたカスタムモデルを学習する方法及びそれを用いた学習装置{method for training custom model based on pre-trained base model and learning device using the same}
CN119418321B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 模型训练方法、用于检测和识别文本的方法及相关装置
CN119418319B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 模型训练方法、文本检测方法、装置、介质和程序产品
CN119418320B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 一种模型训练方法、装置、介质和程序产品
CN119915374B (zh) * 2025-04-03 2025-11-14 浙江潮汐力科技有限公司 故障监测方法、装置、设备、存储介质和程序产品

Family Cites Families (8)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11138392B2 (en) * 2018-07-26 2021-10-05 Google Llc Machine translation using neural network models
US20200042864A1 (en) 2018-08-02 2020-02-06 Veritone, Inc. Neural network orchestration
US11556778B2 (en) 2018-12-07 2023-01-17 Microsoft Technology Licensing, Llc Automated generation of machine learning models
US20210279596A1 (en) 2020-03-06 2021-09-09 Hitachi, Ltd. System for predictive maintenance using trace norm generative adversarial networks
US20220094713A1 (en) * 2020-09-21 2022-03-24 Sophos Limited Malicious message detection
US12141701B2 (en) * 2021-01-21 2024-11-12 International Business Machines Corporation Channel scaling: a scale-and-select approach for selective transfer learning
US11875898B2 (en) * 2021-05-26 2024-01-16 Merative Us L.P. Automatic condition diagnosis using an attention-guided framework
US20230106669A1 (en) * 2021-09-27 2023-04-06 X Development Llc Binding affinity prediction using neural networks

Also Published As

Publication number Publication date
US20260080864A1 (en) 2026-03-19
GB2631139A (en) 2024-12-25
US12512091B2 (en) 2025-12-30
CN118140230A (zh) 2024-06-04
KR20240089615A (ko) 2024-06-20
US20230115321A1 (en) 2023-04-13
GB202403625D0 (en) 2024-04-24
WO2023064033A1 (en) 2023-04-20

Similar Documents

Publication Publication Date Title
JP7561836B2 (ja) 自然言語処理のためのストップワードデータ拡張
JP7703667B2 (ja) 固有表現認識モデルを用いたコンテキストタグ統合
US12299402B2 (en) Techniques for out-of-domain (OOD) detection
JP7721559B2 (ja) 自然言語処理のためのノイズデータ拡張
US12099816B2 (en) Multi-factor modelling for natural language processing
US12512091B2 (en) Fine-tuning multi-head network from a single transformer layer of pre-trained language model
JP2025118956A (ja) 堅牢な固有表現認識のためのチャットボットにおけるエンティティレベルデータ拡張
JP7778160B2 (ja) 単純で効果的な敵対的攻撃方法としてのバリアント不一致攻撃(via)
JP7726995B2 (ja) 自然言語処理のための強化されたロジット
JP7828346B2 (ja) 自然言語処理のためのキーワードデータ拡張ツール
JP2024540111A (ja) 文書からの埋め込まれるデータの抽出のための深層学習技術
JP2024543062A (ja) 自然言語処理のパスのドロップアウト
US20230206125A1 (en) Lexical dropout for natural language processing
JP2024541762A (ja) 事前トレーニングされた言語モデルのための長いテキストを処理するシステムおよび技術
KR20250029146A (ko) 개체-인식 데이터 증강을 위한 기술들
JP2025528391A (ja) 名前付きエンティティ認識モデルの訓練を容易にするための適応的訓練データ拡大

Legal Events

Date Code Title Description
A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20250725

A621 Written request for application examination

Free format text: JAPANESE INTERMEDIATE CODE: A621

Effective date: 20250725