[{"data":1,"prerenderedAt":-1},["ShallowReactive",2],{"doc-detail-85541-en":3,"doc-seo-85541-105":30,"detail-sidebar-cat-0-en-105":91},{"code":4,"msg":5,"data":6},0,"success",{"doc_id":7,"user_id":8,"nickname":9,"user_avatar":10,"doc_module":4,"category_id":11,"category_name":12,"doc_title":13,"doc_description":14,"doc_content":15,"file_id":16,"file_url":17,"file_type":18,"file_size":19,"view_count":20,"is_deleted":4,"is_public":21,"is_downloadable":21,"audit_status":21,"page_count":22,"language":23,"language_code":24,"site_id":25,"html_lang":24,"table_of_contents":26,"faqs":27,"seo_title":13,"seo_description":14,"update_tm":28,"read_time":29},85541,8796095360427,"Lucas Martin","https://ap-avatar.wpscdn.com/davatar_994ba38a5ba835b3df7d355c54d3ed8d",8,"Research & Report","DNA-MGC+ A versatile codec for reliable and resource-efficient data storage on synthetic DNA","DNA-MGC+ addresses the core challenge of DNA data storage where biochemical operations for synthesis, amplification, and sequencing are inherently noisy, generating base-level insertion, deletion, and substitution (IDS) errors and sequence-level dropouts. The codec is evaluated across extensive in silico and in vitro conditions, including Illumina and Nanopore sequencing. Results show consistent improvements over prior codecs, delivering simultaneous gains in sequencing depth, read cost, decoding time, storage density, and error-correction capability. The reliability and efficiency persist on larger files and remain robust beyond benchmark scalability limits.","arXiv :2603 . 14527v2 [ cs .IT] 11 Jul 2026  \nDNA-MGC+: A versatile codec for reliable and resource-efficient data storage on synthetic DNA  \nRamy Khabbaz 1 , Jrmy Mateos 1,2 , Marc Antonini 1,2 , and Serge Kas Hanna 1,*  \n1 Cte d’Azur University, CNRS, I3S, Sophia Antipolis, France  \n2 Pearcode, Sophia Antipolis, France  \n* [serge.kas-hanna@cnrs.fr](serge.kas-hanna@cnrs.fr)  \nABSTRACT  \nThe biochemical processes underlying DNA data storage, including synthesis, amplification, and sequencing, are inherently noisy. Consequently, base-level insertion, deletion, and substitution (IDS) errors, as well as sequence-level dropouts, occur and pose major challenges for reliable data retrieval. Here we introduce DNA-MGC+, a DNA storage codec designed to enable reliable and resource-efficient data retrieval under diverse operating conditions. We evaluate DNA-MGC+ across a wide range of in silico and in vitro settings, including experiments with both Illumina and Nanopore sequencing, and show that it consistently outperforms several representative codecs from the literature. In particular, DNA-MGC+ achieves simultaneous gains in sequencing depth requirements, read cost, decoding time, storage density, and error-correction capability under explicit reliability constraints. These gains persist and become more pronounced for larger files, with DNA-MGC+ remaining reliable and efficient well beyond the practical scalability limits of the benchmarked codecs. Notable performance results include reliable decoding under IDS error rates of up to 24% in synthetic scenarios, and reliable retrieval at sequencing depths below 3× with read costs below 3.5 nts/bit under electrochemical synthesis for both Illumina and Nanopore sequencing.  \nIntroduction  \nThe exponential growth of digital information has made conventional storage technologies increasingly unsustainable 1 , motivating research into alternative storage media, including the use of biological molecules. DNA has emerged as a particularly promising medium owing to its exceptional density and durability2, 3. One gram of DNA can theoretically store exabytes of data, and DNA molecules can remain stable for centuries at room temperature under suitable conditions2, 3 , enabling compact and sustainable long-term storage. The DNA storage pipeline consists of a writing process, a storage phase, and a reading process. During writing, binary information is digitally encoded into quaternary sequences over the four DNA nucleotides Adenine (A), Guanine (G), Cytosine (C), and Thymine (T), which are then synthesized as short DNA molecules known as oligonucleotides. The encoded sequences must be short to adhere to the length limitations imposed by current DNA synthesis technologies. During reading, the stored oligonucleotides are amplified, sequenced, and the resulting reads are digitally processed and decoded to retrieve the original information.  \nWhile several proof-of-concept experiments have already demonstrated the feasibility of this pipeline, scalability remains a central challenge. The main obstacles are related to reliability, speed, and cost. DNA synthesis, amplification, and sequencing are inherently noisy biochemical processes that introduce errors and biases throughout the pipeline4, 5. Existing prototypes predominantly rely on high-fidelity synthesis and sequencing technologies, such as material deposition-based synthesis (Twist Bioscience) and Illumina sequencing, to mitigate these effects. Such technologies, however, are slow and expensive, limiting current systems to small-scale demonstrations that store megabytes of data rather than the exabytes envisioned. Faster and lower-cost alternatives, such as photolithographic synthesis, offer more scalable solutions but exhibit significantly higher error rates6, 7. A key step toward scalability is therefore to address these errors and biases algorithmically rather than preventing them biochemically through high-fidelity technologies, thereby enablin","cbCaipf5FRaTIAhx","https://ap.wps.com/l/cbCaipf5FRaTIAhx","pdf",1530981,2,1,36,"English","en",105,"# Abstract\n# Introduction\n## DNA data storage pipeline and scalability challenges\n## Error mechanisms: base-level IDS and sequence-level dropouts\n## Two-layer error-correcting code architecture\n## Benefits of error-correcting codes in DNA storage","[{\"question\":\"What types of errors does DNA-MGC+ target in DNA data storage?\",\"answer\":\"DNA-MGC+ is designed to handle base-level insertion, deletion, and substitution (IDS) errors as well as sequence-level dropouts that can fully remove encoded sequences from reads.\"},{\"question\":\"How is DNA-MGC+ evaluated and what sequencing platforms are used?\",\"answer\":\"DNA-MGC+ is assessed across in silico and in vitro settings, including experiments using both Illumina and Nanopore sequencing to test performance under diverse operating conditions.\"},{\"question\":\"What performance improvements does DNA-MGC+ achieve compared with existing codecs?\",\"answer\":\"DNA-MGC+ improves reliability while simultaneously reducing sequencing depth requirements, read cost, and decoding time, and it increases storage density and error-correction capability under explicit reliability constraints.\"}]",1784204316,91,{"code":4,"msg":31,"data":32},"ok",{"site_id":25,"language":24,"slug":33,"title":13,"keywords":34,"description":14,"schema_data":35,"social_meta":86,"head_meta":88,"extra_data":90,"updated_unix":28},"dna-mgc-a-versatile-codec-for-reliable-and-resource-efficient-data-storage-on-synthetic-dna","",{"@graph":36,"@context":85},[37,53,68],{"@type":38,"itemListElement":39},"BreadcrumbList",[40,44,47,50],{"item":41,"name":42,"@type":43,"position":21},"https://docshare.wps.com","Home","ListItem",{"item":45,"name":46,"@type":43,"position":20},"https://docshare.wps.com/document/","Document",{"item":48,"name":12,"@type":43,"position":49},"https://docshare.wps.com/document/research-report/",3,{"item":51,"name":13,"@type":43,"position":52},"https://docshare.wps.com/document/dna-mgc-a-versatile-codec-for-reliable-and-resource-efficient-data-storage-on-synthetic-dna/85541/",4,{"url":51,"name":13,"@type":54,"author":55,"headline":13,"publisher":57,"fileFormat":60,"inLanguage":24,"description":14,"dateModified":61,"datePublished":62,"encodingFormat":60,"isAccessibleForFree":63,"interactionStatistic":64},"DigitalDocument",{"name":9,"@type":56},"Person",{"url":41,"name":58,"@type":59},"DocShare","Organization","application/pdf","2026-07-22","2026-07-16",true,{"@type":65,"interactionType":66,"userInteractionCount":20},"InteractionCounter",{"@type":67},"ViewAction",{"@type":69,"mainEntity":70},"FAQPage",[71,77,81],{"name":72,"@type":73,"acceptedAnswer":74},"What types of errors does DNA-MGC+ target in DNA data storage?","Question",{"text":75,"@type":76},"DNA-MGC+ is designed to handle base-level insertion, deletion, and substitution (IDS) errors as well as sequence-level dropouts that can fully remove encoded sequences from reads.","Answer",{"name":78,"@type":73,"acceptedAnswer":79},"How is DNA-MGC+ evaluated and what sequencing platforms are used?",{"text":80,"@type":76},"DNA-MGC+ is assessed across in silico and in vitro settings, including experiments using both Illumina and Nanopore sequencing to test performance under diverse operating conditions.",{"name":82,"@type":73,"acceptedAnswer":83},"What performance improvements does DNA-MGC+ achieve compared with existing codecs?",{"text":84,"@type":76},"DNA-MGC+ improves reliability while simultaneously reducing sequencing depth requirements, read cost, and decoding time, and it increases storage density and error-correction capability under explicit reliability constraints.","https://schema.org",{"og:url":51,"og:type":87,"og:title":13,"og:site_name":58,"og:description":14},"article",{"robots":89,"canonical":51},"index,follow",{"doc_id":7,"site_id":25},{"code":4,"msg":5,"data":92},[93,97,101,105,110,115,120,123,128,131,135],{"id":21,"doc_module":4,"doc_module_name":46,"category_name":94,"show_sort_weight":95,"slug":96},"Story & Novel",90,"story-novel",{"id":20,"doc_module":4,"doc_module_name":46,"category_name":98,"show_sort_weight":99,"slug":100},"Literature",80,"literature",{"id":52,"doc_module":4,"doc_module_name":46,"category_name":102,"show_sort_weight":103,"slug":104},"Exam",70,"exam",{"id":106,"doc_module":4,"doc_module_name":46,"category_name":107,"show_sort_weight":108,"slug":109},5,"Comic",60,"comic",{"id":111,"doc_module":4,"doc_module_name":46,"category_name":112,"show_sort_weight":113,"slug":114},6,"Technology",50,"technology",{"id":116,"doc_module":4,"doc_module_name":46,"category_name":117,"show_sort_weight":118,"slug":119},7,"Healthcare",40,"healthcare",{"id":11,"doc_module":4,"doc_module_name":46,"category_name":12,"show_sort_weight":121,"slug":122},30,"research-report",{"id":124,"doc_module":4,"doc_module_name":46,"category_name":125,"show_sort_weight":126,"slug":127},9,"Religion & Spirituality",20,"religion-spirituality",{"id":126,"doc_module":4,"doc_module_name":46,"category_name":129,"show_sort_weight":126,"slug":130},"World Cup","world-cup",{"id":132,"doc_module":4,"doc_module_name":46,"category_name":133,"show_sort_weight":132,"slug":134},10,"Lifestyle","lifestyle",{"id":136,"doc_module":4,"doc_module_name":46,"category_name":137,"show_sort_weight":106,"slug":138},19,"General","general"]