Reading Nuclear and Chloroplast DNA from the Same Sequencing Reads
With Ms. Eranga Pawani Witharana as the first author, a research group centered in our laboratory demonstrated a method for efficiently recovering nuclear phylogenetic information from the same sequencing reads used for chloroplast genome analysis, enabling an integrated comparison of both data sources.
Original Article Information
Title: Subfamily evolution analysis using nuclear and chloroplast data from the same reads
Authors: Eranga Pawani Witharana, Takaya Iwasaki, Myat Htoo San, Nadeeka U. Jayawardana, Nobuhiro Kotoda, Masashi Yamamoto & Yukio Nagano
Journal: Scientific Reports
Published: January 3, 2025 (accepted December 13, 2024)
DOI: 10.1038/s41598-024-83292-9
Journal metrics: 2025 Journal Impact Factor: 4.9; 2025 5-year Journal Impact Factor: 4.8; 2025 SNIP: 1.339; 2025 SJR: 0.893; the world’s second-most-cited journal (latest publisher-listed metrics as of July 2026)
Article type: Open Access
Chloroplast (cp) genomes are widely used in plant phylogenetics because they are relatively small and tractable. In many angiosperms, however, cpDNA is inherited uniparentally—often maternally—and can therefore carry a different phylogenetic signal from biparentally inherited nuclear DNA. Chloroplast data alone may consequently be insufficient to resolve some evolutionary relationships.
To obtain more comprehensive phylogenetic information, nuclear DNA information, which is inherited from both parents, is also important. Nuclear DNA provides a more holistic genetic perspective than cpDNA. However, acquiring full nuclear DNA information was previously associated with high costs. In recent years, highly cost-effective methods for obtaining nuclear DNA data have emerged (e.g., target capture, RNA sequencing), but it has generally been necessary to prepare chloroplast and nuclear DNA data separately.
In this study, to address this challenge, we focused on the development and validation of a new method that obtains nuclear DNA phylogenetic information from the “exact same” raw (unprocessed) read sequence data used for chloroplast genome analysis.
We utilized a new computational tool called Read2Tree. Read2Tree enables the direct and efficient extraction of conserved nuclear gene sequences from raw read sequence data.
Following publication of the paper, we released the analysis pipeline on GitHub at “Eranga-Witharana/rt2-phylogenomics” to facilitate reproduction and further application of the workflow.
Main Features and Advantages of Read2Tree
- Cost-effective: Using the same raw data as chloroplast analysis reduces data acquisition costs.
- Comprehensive nuclear data retrieval: A reference genome-guided approach allows extraction of a large number of conserved nuclear genes.
- Streamlined analysis: Functions as an integrated pipeline from gene identification through alignment generation to phylogenetic tree construction.
- No custom preparation required: Uses existing reference datasets, eliminating the need to design and validate species-specific probe sets or databases.
- Useful for discordance analysis: Outputs individual gene alignment data, which can be used to investigate causes of phylogenetic discordance (such as discrepancies between gene trees and species trees).
Method Validation: An Example Using Aurantioideae Plants
To validate the effectiveness and usefulness of this new method (nuclear DNA phylogenetic analysis using Read2Tree), we targeted the plant group Aurantioideae. Aurantioideae, which includes citrus and its close relatives, is an important group with complex phylogenetic relationships.
Using Read2Tree, the study recovered conserved nuclear gene sequences for 39 Aurantioideae species and compared the resulting nuclear phylogenies with chloroplast genome trees built from the same raw reads.
- Comparison with existing nuclear DNA methods: The nuclear DNA phylogenetic tree obtained by Read2Tree was compared with the tree generated by RAD-Seq, another nuclear DNA method used in previous studies. This comparison evaluated the reliability of Read2Tree.
- Integration with chloroplast phylogenetic analysis: Phylogenetic trees of chloroplast genomes were constructed from the same raw data and compared with the nuclear DNA trees produced by Read2Tree. This approach investigated cytonuclear discordance, the mismatch between chloroplast and nuclear phylogenies.
- Analysis of discordance causes: Using individual gene alignment data output by Read2Tree and combining it with other methods (D-statistic, DFOIL, QuIBL, etc.), we analyzed whether observed phylogenetic discordance was caused by complex evolutionary processes such as incomplete lineage sorting (ILS), introgression, or ancient introgression.
Results (Method Effectiveness)
The results of this study revealed the following points, demonstrating the effectiveness of the method using Read2Tree:
- Comparison of nuclear phylogenies: The Read2Tree nuclear phylogenies were largely concordant with relationships reported by an earlier RAD-Seq study. Some taxon placements remained uncertain, so the analysis should not be read as resolving every branch definitively.
- Efficient reuse of data: The same short-read dataset supported both chloroplast genome analysis and nuclear phylogenetic inference, making the approach more cost-efficient than preparing separate sequencing datasets. This was a data-acquisition advantage, not a formal economic evaluation.
- Analysis of phylogenetic discordance: Multiple analyses suggested that incomplete lineage sorting (ILS) explains most of the observed discordance, while introgression and ancient introgression may also have contributed within particular clades.
Conclusion
The study showed that Read2Tree can recover conserved nuclear genes from the same raw reads used for chloroplast genome analysis, enabling comparison of nuclear and chloroplast phylogenies and investigation of contributions from incomplete lineage sorting (ILS) and introgression. Some taxon placements remained uncertain, and the study did not resolve every branch definitively.
This method has the potential to contribute to more detailed and reliable reconstructions of evolutionary histories in future plant phylogenetic research.
Potential Applications to Research on Food Ingredients
The following points describe potential applications inferred from the validated phylogenetic workflow. The paper did not directly test crop breeding, flavor, nutritional value, or disease resistance.
For example, groups like the Aurantioideae used in this study include plant species that are important for food and medicinal uses worldwide. Understanding their precise phylogenetic relationships and the complex evolutionary history of hybridization is crucial for crop improvement in agriculture and for the conservation and utilization of valuable genetic resources (such as wild relatives) that may become future food sources.
This new method using Read2Tree enables efficient retrieval of information from both chloroplast and nuclear DNA using the same dataset, allowing more comprehensive and detailed phylogenetic analyses of many food crop species that were previously difficult to analyze.
Moreover, the ability to analyze complex genetic processes such as incomplete lineage sorting (ILS) and introgression provides a foundation for understanding, at the genetic level, how diverse traits of food plants—such as high nutritional value, specific flavors, and resistance to pests and diseases—have been acquired and maintained throughout evolutionary history.
In this way, this method can be said to offer an important foundation for deepening our understanding of the diversity of plants used as food and for driving research toward the development of improved varieties and sustainable resource utilization.