GB2631139A - Fine-tuning multi-head network from a single transformer layer of pre-trained language model - Google Patents

Fine-tuning multi-head network from a single transformer layer of pre-trained language model Download PDF

Info

Publication number
GB2631139A
GB2631139A GB2403625.3A GB202403625A GB2631139A GB 2631139 A GB2631139 A GB 2631139A GB 202403625 A GB202403625 A GB 202403625A GB 2631139 A GB2631139 A GB 2631139A
Authority
GB
United Kingdom
Prior art keywords
multiple layers
learning model
machine
parameter values
layer
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Pending
Application number
GB2403625.3A
Other languages
English (en)
Other versions
GB202403625D0 (en
Inventor
Tien Vu Thanh
Quang Pham Tuyen
Mohamad Nezami Omid
Edward Johnson Mark
Long Duong Thanh
Duy Vu Hoang Cong
Current Assignee (The listed assignees may be inaccurate. Google has not performed a legal analysis and makes no representation or warranty as to the accuracy of the list.)
Oracle International Corp
Original Assignee
Oracle International Corp
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by Oracle International Corp filed Critical Oracle International Corp
Publication of GB202403625D0 publication Critical patent/GB202403625D0/en
Publication of GB2631139A publication Critical patent/GB2631139A/en
Pending legal-status Critical Current

Links

Classifications

    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/20Natural language analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06FELECTRIC DIGITAL DATA PROCESSING
    • G06F40/00Handling natural language data
    • G06F40/30Semantic analysis
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N20/00Machine learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/045Combinations of networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/04Architecture, e.g. interconnection topology
    • G06N3/0475Generative networks
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/084Backpropagation, e.g. using gradient descent
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/09Supervised learning
    • GPHYSICS
    • G06COMPUTING OR CALCULATING; COUNTING
    • G06NCOMPUTING ARRANGEMENTS BASED ON SPECIFIC COMPUTATIONAL MODELS
    • G06N3/00Computing arrangements based on biological models
    • G06N3/02Neural networks
    • G06N3/08Learning methods
    • G06N3/096Transfer learning
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • HELECTRICITY
    • H04ELECTRIC COMMUNICATION TECHNIQUE
    • H04LTRANSMISSION OF DIGITAL INFORMATION, e.g. TELEGRAPHIC COMMUNICATION
    • H04L51/00User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail
    • H04L51/02User-to-user messaging in packet-switching networks, transmitted according to store-and-forward or real-time protocols, e.g. e-mail using automatic reactions or user delegation, e.g. automatic replies or chatbot-generated messages
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/06Creation of reference templates; Training of speech recognition systems, e.g. adaptation to the characteristics of the speaker's voice
    • G10L15/063Training
    • G10L2015/0635Training updating or merging of old and new templates; Mean values; Weighting
    • GPHYSICS
    • G10MUSICAL INSTRUMENTS; ACOUSTICS
    • G10LSPEECH ANALYSIS TECHNIQUES OR SPEECH SYNTHESIS; SPEECH RECOGNITION; SPEECH OR VOICE PROCESSING TECHNIQUES; SPEECH OR AUDIO CODING OR DECODING
    • G10L15/00Speech recognition
    • G10L15/22Procedures used during a speech recognition process, e.g. man-machine dialogue
    • G10L2015/223Execution procedure of a spoken command

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Physics & Mathematics (AREA)
  • Health & Medical Sciences (AREA)
  • Computational Linguistics (AREA)
  • Artificial Intelligence (AREA)
  • General Physics & Mathematics (AREA)
  • General Engineering & Computer Science (AREA)
  • General Health & Medical Sciences (AREA)
  • Software Systems (AREA)
  • Computing Systems (AREA)
  • Mathematical Physics (AREA)
  • Data Mining & Analysis (AREA)
  • Evolutionary Computation (AREA)
  • Biomedical Technology (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Molecular Biology (AREA)
  • Biophysics (AREA)
  • Audiology, Speech & Language Pathology (AREA)
  • Acoustics & Sound (AREA)
  • Human Computer Interaction (AREA)
  • Multimedia (AREA)
  • Computer Networks & Wireless Communication (AREA)
  • Signal Processing (AREA)
  • Medical Informatics (AREA)
  • Computer Vision & Pattern Recognition (AREA)
  • Machine Translation (AREA)
  • Information Retrieval, Db Structures And Fs Structures Therefor (AREA)
GB2403625.3A 2021-10-12 2022-08-17 Fine-tuning multi-head network from a single transformer layer of pre-trained language model Pending GB2631139A (en)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US202163254740P 2021-10-12 2021-10-12
US17/735,651 US12512091B2 (en) 2021-10-12 2022-05-03 Fine-tuning multi-head network from a single transformer layer of pre-trained language model
PCT/US2022/040530 WO2023064033A1 (en) 2021-10-12 2022-08-17 Fine-tuning multi-head network from a single transformer layer of pre-trained language model

Publications (2)

Publication Number Publication Date
GB202403625D0 GB202403625D0 (en) 2024-04-24
GB2631139A true GB2631139A (en) 2024-12-25

Family

ID=85798249

Family Applications (1)

Application Number Title Priority Date Filing Date
GB2403625.3A Pending GB2631139A (en) 2021-10-12 2022-08-17 Fine-tuning multi-head network from a single transformer layer of pre-trained language model

Country Status (6)

Country Link
US (2) US12512091B2 (https=)
JP (1) JP2024539003A (https=)
KR (1) KR20240089615A (https=)
CN (1) CN118140230A (https=)
GB (1) GB2631139A (https=)
WO (1) WO2023064033A1 (https=)

Families Citing this family (15)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US12548552B2 (en) * 2021-11-19 2026-02-10 International Business Machines Corporation Dynamic language selection of an AI voice assistance system
US11947935B2 (en) * 2021-11-24 2024-04-02 Microsoft Technology Licensing, Llc. Custom models for source code generation via prefix-tuning
US20240061835A1 (en) * 2022-08-22 2024-02-22 Oracle International Corporation System and method of selective fine-tuning for custom training of a natural language to logical form model
US20240169165A1 (en) * 2022-11-17 2024-05-23 Samsung Electronics Co., Ltd. Automatically Generating Annotated Ground-Truth Corpus for Training NLU Model
US12562163B2 (en) * 2023-05-12 2026-02-24 Servicenow, Inc. Bidirectional assistant for development platforms
CN116774140A (zh) * 2023-06-26 2023-09-19 南京邮电大学 基于残差注意力网络的无网格信号源doa估计方法
US20250005282A1 (en) * 2023-06-29 2025-01-02 Amazon Technologies, Inc. Domain entity extraction for performing text analysis tasks
CN118446218B (zh) * 2024-05-16 2024-11-01 西南交通大学 一种对抗式阅读理解嵌套命名实体识别方法
CA3253531A1 (en) * 2024-06-14 2026-01-19 The Toronto-Dominion Bank Context retrieval for in-context learning model
WO2026000314A1 (en) * 2024-06-27 2026-01-02 Beijing Youzhuju Network Technology Co., Ltd. Model-based task processing
JP7658644B1 (ja) * 2024-10-21 2025-04-08 スパーブエーアイ カンパニー リミテッド 事前学習されたベースモデルに基づいたカスタムモデルを学習する方法及びそれを用いた学習装置{method for training custom model based on pre-trained base model and learning device using the same}
CN119418321B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 模型训练方法、用于检测和识别文本的方法及相关装置
CN119418319B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 模型训练方法、文本检测方法、装置、介质和程序产品
CN119418320B (zh) * 2024-10-30 2025-09-30 上海哔哩哔哩科技有限公司 一种模型训练方法、装置、介质和程序产品
CN119915374B (zh) * 2025-04-03 2025-11-14 浙江潮汐力科技有限公司 故障监测方法、装置、设备、存储介质和程序产品

Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20200042864A1 (en) * 2018-08-02 2020-02-06 Veritone, Inc. Neural network orchestration
US20200184327A1 (en) * 2018-12-07 2020-06-11 Microsoft Technology Licensing, Llc Automated generation of machine learning models
US20210279596A1 (en) * 2020-03-06 2021-09-09 Hitachi, Ltd. System for predictive maintenance using trace norm generative adversarial networks
US11138392B2 (en) * 2018-07-26 2021-10-05 Google Llc Machine translation using neural network models

Family Cites Families (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US20220094713A1 (en) * 2020-09-21 2022-03-24 Sophos Limited Malicious message detection
US12141701B2 (en) * 2021-01-21 2024-11-12 International Business Machines Corporation Channel scaling: a scale-and-select approach for selective transfer learning
US11875898B2 (en) * 2021-05-26 2024-01-16 Merative Us L.P. Automatic condition diagnosis using an attention-guided framework
US20230106669A1 (en) * 2021-09-27 2023-04-06 X Development Llc Binding affinity prediction using neural networks

Patent Citations (4)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US11138392B2 (en) * 2018-07-26 2021-10-05 Google Llc Machine translation using neural network models
US20200042864A1 (en) * 2018-08-02 2020-02-06 Veritone, Inc. Neural network orchestration
US20200184327A1 (en) * 2018-12-07 2020-06-11 Microsoft Technology Licensing, Llc Automated generation of machine learning models
US20210279596A1 (en) * 2020-03-06 2021-09-09 Hitachi, Ltd. System for predictive maintenance using trace norm generative adversarial networks

Non-Patent Citations (1)

* Cited by examiner, † Cited by third party
Title
YUNHUI GUO et al.,"SpotTune: Transfer Learning Through Adaptive Fine-Tuning", 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 09 January 2020, pages 4800-4807; and figures 1-2 *

Also Published As

Publication number Publication date
US20260080864A1 (en) 2026-03-19
JP2024539003A (ja) 2024-10-28
US12512091B2 (en) 2025-12-30
CN118140230A (zh) 2024-06-04
KR20240089615A (ko) 2024-06-20
US20230115321A1 (en) 2023-04-13
GB202403625D0 (en) 2024-04-24
WO2023064033A1 (en) 2023-04-20

Similar Documents

Publication Publication Date Title
GB2631139A (en) Fine-tuning multi-head network from a single transformer layer of pre-trained language model
US10978052B2 (en) Email-like user interface for training natural language systems
IL307851B1 (en) Dynamic response prediction for improved robot task processing
AU2022200432B2 (en) Extensible search, content, and dialog management system
MX2024006099A (es) Tecnicas de ajuste basadas en logica de aprendizaje automatico para robots.
JPWO2023064033A5 (https=)
GB2622755A (en) Evaluating output sequences using an auto-regressive language model neural network
JPWO2021050170A5 (https=)
GB2625476A (en) Path dropout for natural language processing
DE102016125954A1 (de) Sprachwiedererkennung mit externen Datenquellen
GB2602920A (en) Deep learning seismic attribute fault predictions
WO2021108796A3 (en) System and method of federated learning with diversified feedback
DE102016125823B4 (de) Unterstützung bei der semantischen offline-bearbeitung durch ein gerät mit begrenzten möglichkeiten
GB2604276A (en) Rare topic detection using hierarchical clustering
MX2020002519A (es) Metodo implementado por computadora para el monitoreo de varias maquinas de procesamiento de cables y sistema de monitoreo.
CN113094467A (zh) 一种知识图谱的查询方法、电子设备及存储介质
GB2612275A (en) Drilling data correction with machine learning and rules-based predictions
CO2022014596A2 (es) Sistema, método y producto de programa de computadora para optimizar un proceso de manufactura
JP2025024163A (ja) 画像編集方法、装置、電子機器及び記憶媒体
WO2021118532A8 (en) Systems and methods for interpreting a voice query
CN111460303A (zh) 数据处理方法、装置、电子设备及计算机可读存储介质
FI3607436T3 (fi) Latenttien aiheuttajien erittely tietokonejärjestelmän optimoimiseksi
MacAvaney et al. A reproducibility study of plaid
CN209657301U (zh) 一种智能翻译机
BR112023019971A2 (pt) Reconhecimento de fala visual adaptativo