arXiv ScienceSearch

arXiv · 2412.06521

Ancient DNA from 120-Million-Year-Old Lycoptera Fossils Reveals Evolutionary Insights

Abstract

High quality ancient DNA (aDNA) is essential for molecular paleontology. Due to DNA degradation and contamination by environmental DNA (eDNA), current research is limited to fossils less than 1 million years old. The study successfully extracted DNA from Lycoptera davidi fossils from the Early Cretaceous period, dating 120 million years ago. Using high-throughput sequencing, 1,258,901 DNA sequences were obtained. We established a rigorous protocol known as the mega screen method. Using this method, we identified 243 original in situ DNA (oriDNA) sequences, likely from the Lycoptera genome. These sequences have an average length of over 100 base pairs and show no signs of deamination. Additionally, 10 transposase coding sequences were discovered, shedding light on a unique self-renewal mechanism in the genome. This study provides valuable DNA data for understanding ancient fish evolution and advances paleontological research.

Explore related subjects

Keep this discovery

Explore connections, maps & timelines

BibTeXRIS

Wan-Qian Zhao, Zhan-Yong Guo, Zeng-Yuan Tian, Tong-Fu Su, Gang-Qiang Cao, Zi-Xin Qi, Tian-Cang Qin, Wei Zhou, Jin-Yu Yang, Ming-Jie Chen, Xin-Ge Zhang, Chun-Yan Zhou, Chuan-Jia Zhu, Meng-Fei Tang, Di Wu, Mei-Rong Song, Yu-Qi Guo, Li-You Qiu, Fei Liang, Mei-Jun Li, Jun-Hui Geng, Li-Juan Zhao, Shu-Jie Zhang. 2024-12-09. Ancient DNA from 120-Million-Year-Old Lycoptera Fossils Reveals Evolutionary Insights. https://arxiv.org/abs/2412.06521

Cite the original work for its findings. Save a collection to share your selection of sources.

KEEP EXPLORING

Related papers

PlainMap: a lightweight, restartable mapping pipeline for ancient and modern DNA

Mapping sequencing reads to a reference genome requires preprocessing and alignment choices that can vary with library type, fragment length, sequencing platform, and reference genome. These considerations are particularly important for ancient and historical DNA, where short and damaged fragments can make the most appropriate mapping strategy difficult to determine a priori. We present PlainMap, a lightweight and restartable mapping pipeline for modern and degraded DNA sequencing data. PlainMap accepts a simple manifest of FASTQ files, automatically identifies single-end and paired-end data from read headers, supports mixed sequencing platforms, and provides alternative mapping strategies for modern and degraded DNA. Deterministic chunking and checkpoint-based execution allow large analyses to resume after interruption, while optional pilot subsampling enables empirical comparison of mapping strategies using identical subsets of raw fragments. PlainMap produces duplicate-filtered BAM files together with fragment-aware mapping and coverage statistics. Evaluation using heterogeneous sequencing data confirmed the expected behaviour of the three mapping modes, while adaptive chunking reduced peak memory use by approximately 25% and allowed an interrupted analysis to resume from completed mapping chunks. PlainMap is implemented as a single Bash script and is freely available at https://github.com/BiodiversityExtinction/PlainMap.

q-bio.GN

Democratizing Clinical Tumor Whole Genome Sequencing: 18-hour End-to-end Analysis via Trillion-parameter Large Language Models Locally Deployed on Consumer-grade Hardware

Whole genome sequencing (WGS) is essential for precision oncology, yet its clinical adoption remains limited by prohibitive computational costs and multi-day turnaround times. This work presents a fully localized low-resource framework enabling stable deployment of a trillion-parameter biomedical LLM on a single consumer-grade RTX 4060 laptop with 32GB system memory and 8GB VRAM, as well as on routine clinical workstations in general hospitals, completing the entire tumor-paired WGS workflow from raw FASTQ input to clinical-grade full-variation-spectrum report output. Under standard 30X depth configurations, our implementation finishes a single tumor-paired WGS analysis within 18 hours, achieving 99.62% F1 score for somatic variant detection with over 99.9% concordance to the industrial-standard A100 cluster pipeline, fully meeting clinical oncology accuracy requirements. Quantitative profiling shows adaptive heterogeneous memory scheduling accounts for 71% of total execution time, while model optimization introduces less than 9% of total detection error. This work is the first engineering implementation of trillion-parameter biomedical LLM-driven clinical-grade genomic analysis on consumer-grade hardware, breaking the industry paradigm that trillion-scale genomic LLMs require hundred-thousand-dollar GPU clusters and multi-day turnaround, establishing a low-resource pathway for global primary medical institutions to adopt whole-genome precision oncology at zero additional cost.

q-bio.GN

Structure is not mechanism: high-gain gated-FFN rows across text and genomic foundation models

A small number of unusually high-gain parameters can exert disproportionate effects in transformer language models, but whether analogous structures recur in genomic foundation models and whether structural geometry determines functional importance remains unknown. We analyzed high-gain rows in gated feed-forward networks across text and genomic foundation models, including a frozen 22-model causal census. Computing an associated bilinear weight operator exactly, without a diagonal approximation, we tested whether structural extremeness is a transferable mechanism. Activation-derived candidates were functionally enriched relative to random and top-norm same-layer controls, yet neither spectral concentration nor operator magnitude predicted causal effect size, and these associations vanished within the endpoint-homogeneous text-decoder subset. A within-layer sweep of 36 rows in one genomic and one text decoder resolved this into two regimes: below the detector's acceptance threshold the ratio carried no positive information about causal damage, whereas above it the ratio ordered rows strongly but did not grade severity as a dose-response. The same sweep revealed a second individually catastrophic row invisible to a one-candidate-per-model census, and non-additive damage among co-located critical rows. Case studies showed divergent causal organizations: a robust super-additive pair interaction in DNABERT-2, and in GENERator a sharply position-localized dependence in which preserving or restoring the row's beginning-of-sequence contribution rescued essentially all native-loss damage. High-gain gated-FFN rows are therefore a recurrent architectural phenotype whose structural prominence acts as an enrichment signal, not a calibrated measure of functional criticality or a specification of causal organization. Enrichment is general, but the mechanism is model-specific.

q-bio.GN