JP7810664B2 - 質スコア圧縮 - Google Patents

質スコア圧縮

Info

Publication number
JP7810664B2
JP7810664B2 JP2022575435A JP2022575435A JP7810664B2 JP 7810664 B2 JP7810664 B2 JP 7810664B2 JP 2022575435 A JP2022575435 A JP 2022575435A JP 2022575435 A JP2022575435 A JP 2022575435A JP 7810664 B2 JP7810664 B2 JP 7810664B2
Authority
JP
Japan
Prior art keywords
sequence
encoding
quality score
read
nucleic acid
Prior art date
Legal status (The legal status is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the status listed.)
Active
Application number
JP2022575435A
Other languages
English (en)
Japanese (ja)
Other versions
JP2023547973A (ja
JP2023547973A5 (https=
Inventor
ギヨーム・アレクサンドル・パスカル・リツク
Original Assignee
イルミナ インコーポレイテッド
Priority date (The priority date is an assumption and is not a legal conclusion. Google has not performed a legal analysis and makes no representation as to the accuracy of the date listed.)
Filing date
Publication date
Application filed by イルミナ インコーポレイテッド filed Critical イルミナ インコーポレイテッド
Publication of JP2023547973A publication Critical patent/JP2023547973A/ja
Publication of JP2023547973A5 publication Critical patent/JP2023547973A5/ja
Application granted granted Critical
Publication of JP7810664B2 publication Critical patent/JP7810664B2/ja
Active legal-status Critical Current
Anticipated expiration legal-status Critical

Links

Classifications

    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B30/00ICT specially adapted for sequence analysis involving nucleotides or amino acids
    • GPHYSICS
    • G16INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR SPECIFIC APPLICATION FIELDS
    • G16BBIOINFORMATICS, i.e. INFORMATION AND COMMUNICATION TECHNOLOGY [ICT] SPECIALLY ADAPTED FOR GENETIC OR PROTEIN-RELATED DATA PROCESSING IN COMPUTATIONAL MOLECULAR BIOLOGY
    • G16B50/00ICT programming tools or database systems specially adapted for bioinformatics
    • G16B50/50Compression of genetic data
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/3068Precoding preceding compression, e.g. Burrows-Wheeler transformation
    • H03M7/3071Prediction
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/3084Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction using adaptive string matching, e.g. the Lempel-Ziv method
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/40Conversion to or from variable length codes, e.g. Shannon-Fano code, Huffman code, Morse code
    • H03M7/4006Conversion to or from arithmetic code
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/60General implementation details not specific to a particular type of compression
    • H03M7/6011Encoder aspects
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/60General implementation details not specific to a particular type of compression
    • H03M7/6017Methods or arrangements to increase the throughput
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/60General implementation details not specific to a particular type of compression
    • H03M7/6017Methods or arrangements to increase the throughput
    • H03M7/6029Pipelining
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/70Type of the data to be coded, other than image and sound
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/70Type of the data to be coded, other than image and sound
    • H03M7/702Software
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/70Type of the data to be coded, other than image and sound
    • H03M7/705Unicode
    • HELECTRICITY
    • H03ELECTRONIC CIRCUITRY
    • H03MCODING; DECODING; CODE CONVERSION IN GENERAL
    • H03M7/00Conversion of a code where information is represented by a given sequence or number of digits to a code where the same, similar or subset of information is represented by a different sequence or number of digits
    • H03M7/30Compression; Expansion; Suppression of unnecessary data, e.g. redundancy reduction
    • H03M7/70Type of the data to be coded, other than image and sound
    • H03M7/707Structured documents, e.g. XML

Landscapes

  • Engineering & Computer Science (AREA)
  • Theoretical Computer Science (AREA)
  • Health & Medical Sciences (AREA)
  • Physics & Mathematics (AREA)
  • Life Sciences & Earth Sciences (AREA)
  • Biotechnology (AREA)
  • Medical Informatics (AREA)
  • Spectroscopy & Molecular Physics (AREA)
  • Bioinformatics & Cheminformatics (AREA)
  • Bioinformatics & Computational Biology (AREA)
  • Biophysics (AREA)
  • Evolutionary Biology (AREA)
  • General Health & Medical Sciences (AREA)
  • Bioethics (AREA)
  • Databases & Information Systems (AREA)
  • Genetics & Genomics (AREA)
  • Chemical & Material Sciences (AREA)
  • Analytical Chemistry (AREA)
  • Proteomics, Peptides & Aminoacids (AREA)
  • Measuring Or Testing Involving Enzymes Or Micro-Organisms (AREA)
  • Compression, Expansion, Code Conversion, And Decoders (AREA)
  • Applications Or Details Of Rotary Compressors (AREA)
JP2022575435A 2020-11-05 2021-11-05 質スコア圧縮 Active JP7810664B2 (ja)

Applications Claiming Priority (3)

Application Number Priority Date Filing Date Title
US202063110308P 2020-11-05 2020-11-05
US63/110,308 2020-11-05
PCT/US2021/058364 WO2022099097A1 (en) 2020-11-05 2021-11-05 Quality score compression

Publications (3)

Publication Number Publication Date
JP2023547973A JP2023547973A (ja) 2023-11-15
JP2023547973A5 JP2023547973A5 (https=) 2024-11-13
JP7810664B2 true JP7810664B2 (ja) 2026-02-03

Family

ID=78725748

Family Applications (1)

Application Number Title Priority Date Filing Date
JP2022575435A Active JP7810664B2 (ja) 2020-11-05 2021-11-05 質スコア圧縮

Country Status (12)

Country Link
US (4) US11527307B2 (https=)
EP (1) EP4241276A1 (https=)
JP (1) JP7810664B2 (https=)
KR (1) KR20230101760A (https=)
CN (1) CN115668384A (https=)
AU (1) AU2021376411A1 (https=)
BR (1) BR112022025042A2 (https=)
CA (1) CA3174208A1 (https=)
IL (2) IL298981B2 (https=)
MX (1) MX2022016020A (https=)
WO (1) WO2022099097A1 (https=)
ZA (2) ZA202304367B (https=)

Families Citing this family (5)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2022099097A1 (en) * 2020-11-05 2022-05-12 Illumina, Inc. Quality score compression
JP2022086403A (ja) * 2020-11-30 2022-06-09 キオクシア株式会社 メモリシステム及び情報処理システム
EP4490735A1 (en) 2022-03-08 2025-01-15 Illumina Inc Multi-pass software-accelerated genomic read mapping engine
US11775172B1 (en) * 2022-05-05 2023-10-03 CELLGENTEK Corp. Genome data compression and transmission method for FASTQ-formatted genome data
CN115662525B (zh) * 2022-10-25 2026-04-21 湖南大学 一种测序fastq文件质量分数序列的稀疏化处理方法

Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008108297A1 (ja) 2007-03-02 2008-09-12 Research Organization Of Information And Systems 相同性検索システム
WO2018068830A1 (en) 2016-10-11 2018-04-19 Genomsys Sa Method and system for the transmission of bioinformatics data

Family Cites Families (12)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
US10090857B2 (en) * 2010-04-26 2018-10-02 Samsung Electronics Co., Ltd. Method and apparatus for compressing genetic data
US20110288785A1 (en) * 2010-05-18 2011-11-24 Translational Genomics Research Institute (Tgen) Compression of genomic base and annotation data
AU2012272161B2 (en) * 2011-06-21 2015-12-24 Illumina Cambridge Limited Methods and systems for data analysis
US10777301B2 (en) * 2012-07-13 2020-09-15 Pacific Biosciences For California, Inc. Hierarchical genome assembly method using single long insert library
US10847251B2 (en) * 2013-01-17 2020-11-24 Illumina, Inc. Genomic infrastructure for on-site or cloud-based DNA and RNA processing and analysis
WO2014197377A2 (en) * 2013-06-03 2014-12-11 Good Start Genetics, Inc. Methods and systems for storing sequence read data
WO2016081712A1 (en) * 2014-11-19 2016-05-26 Bigdatabio, Llc Systems and methods for genomic manipulations and analysis
CN107851118A (zh) * 2015-05-21 2018-03-27 基因福米卡数据系统有限公司 下一代测序数据的存储、传输和压缩
CN110021349B (zh) * 2017-07-31 2021-02-02 北京哲源科技有限责任公司 基因数据的编码方法
CN110111852A (zh) * 2018-01-11 2019-08-09 广州明领基因科技有限公司 一种海量dna测序数据无损快速压缩平台
CN110797082A (zh) * 2019-10-24 2020-02-14 福建和瑞基因科技有限公司 基因测序数据的存储读取方法及系统
WO2022099097A1 (en) 2020-11-05 2022-05-12 Illumina, Inc. Quality score compression

Patent Citations (2)

* Cited by examiner, † Cited by third party
Publication number Priority date Publication date Assignee Title
WO2008108297A1 (ja) 2007-03-02 2008-09-12 Research Organization Of Information And Systems 相同性検索システム
WO2018068830A1 (en) 2016-10-11 2018-04-19 Genomsys Sa Method and system for the transmission of bioinformatics data

Also Published As

Publication number Publication date
CN115668384A (zh) 2023-01-31
IL298981B2 (en) 2025-03-01
ZA202304367B (en) 2023-12-20
KR20230101760A (ko) 2023-07-06
IL316156B1 (en) 2026-04-01
IL298981A (en) 2023-02-01
WO2022099097A1 (en) 2022-05-12
ZA202402955B (en) 2025-04-30
BR112022025042A2 (pt) 2023-05-09
IL316156A (en) 2024-12-01
US20220139502A1 (en) 2022-05-05
US20240420804A1 (en) 2024-12-19
JP2023547973A (ja) 2023-11-15
MX2022016020A (es) 2023-02-02
US20230040143A1 (en) 2023-02-09
US12080385B2 (en) 2024-09-03
AU2021376411A1 (en) 2022-10-27
US20240062853A1 (en) 2024-02-22
CA3174208A1 (en) 2022-05-12
US11527307B2 (en) 2022-12-13
IL298981B1 (en) 2024-11-01
EP4241276A1 (en) 2023-09-13
US11776663B2 (en) 2023-10-03

Similar Documents

Publication Publication Date Title
JP7810664B2 (ja) 質スコア圧縮
US10902937B2 (en) Lossless compression of DNA sequences
Wandelt et al. Trends in genome compression
CN113826168B (zh) 用于散列表基因组映射的灵活种子延伸
KR20220034082A (ko) 적대적 네트워크 모델을 트레이닝하는 방법 및 장치, 문자 라이브러리를 구축하는 방법 및 장치, 전자장비, 저장매체 및 컴퓨터 프로그램
US12125562B2 (en) Quality value compression framework in aligned sequencing data based on novel contexts
Li et al. DNA-COMPACT: DNA COM pression Based on a P attern-A ware C ontextual Modeling T echnique
KR20120137235A (ko) 유전자 데이터를 압축하는 방법 및 장치
JP2025124637A (ja) ゲノム配列データの圧縮のための方法
JP2023503739A (ja) 遺伝子融合の迅速な検出
Mansouri et al. One-bit dna compression algorithm
EP3583249A1 (en) Method and systems for the reconstruction of genomic reference sequences from compressed genomic sequence reads
US10460829B2 (en) Systems and methods for encoding genetic variation for a population
CN110377822A (zh) 用于网络表征学习的方法、装置及电子设备
Nazari et al. Lossless and reference-free compression of FASTQ/A files using GeneSqueeze
Sun et al. An intelligent ubiquitous compression technique for DNA sequencing using Hadoop
CN117493467A (zh) 数据模型的主键查找方法、装置、电子设备及存储介质
WO2025179458A1 (zh) 基因测序数据的压缩方法、解压方法及装置
KR20230069046A (ko) 소프트웨어 가속 게놈 판독 매핑
CN112464011A (zh) 数据检索方法及装置

Legal Events

Date Code Title Description
A521 Request for written amendment filed

Free format text: JAPANESE INTERMEDIATE CODE: A523

Effective date: 20241105

A621 Written request for application examination

Free format text: JAPANESE INTERMEDIATE CODE: A621

Effective date: 20241105

TRDD Decision of grant or rejection written
A01 Written decision to grant a patent or to grant a registration (utility model)

Free format text: JAPANESE INTERMEDIATE CODE: A01

Effective date: 20251223

A61 First payment of annual fees (during grant procedure)

Free format text: JAPANESE INTERMEDIATE CODE: A61

Effective date: 20260122

R150 Certificate of patent or registration of utility model

Ref document number: 7810664

Country of ref document: JP

Free format text: JAPANESE INTERMEDIATE CODE: R150