High-Throughput Pipeline for Genome Analysis of Bacteria Used in Food Production
A Bacillus subtilis strain isolated from the traditional Indian fermented food bekang is phylogenetically distinct from Japanese natto strains yet shares a closely related set of accessory genes. This study also presented an efficient comparative-genomics pipeline for bacteria used in food production.
Original Research Information
Authors: Kiyohiko Seki & Yukio Nagano
Journal: Scientific Reports
Published: December 17, 2025 (accepted November 18, 2025)
DOI: 10.1038/s41598-025-29683-y
Journal metrics: 2025 Journal Impact Factor: 4.9; 2025 5-year Journal Impact Factor: 4.8; 2025 SNIP: 1.339; 2025 SJR: 0.893; the world’s second-most-cited journal (latest publisher-listed metrics as of July 2026)
Article type: Open Access
To examine diversity in Bacillus subtilis used in food fermentation, the study initially collected 55 genomes and used 42 quality-controlled, consistently annotated strains for its principal pangenome and comparative-genomics analyses.
Background and Main Findings
For a general overview and the story behind this discovery, please refer to the following links:
- Press Release: Japanese “Natto” and Indian Traditional Food “Bekang” Bacteria: Genetically “Strangers” yet “Identical Twins”!?
- Research Story: The Global "Natto" Connection: How a Bowl of Beans in Myanmar Sparked a Genomic Discovery (Springer Nature Research Communities)
The principal analysis defined 26 Japanese natto isolates as the narrow-sense natto group. One isolate from the Indian traditional food bekang (Food_IND_1) lay outside the natto clade in the core-genome phylogeny but was its closest functional neighbor by Jaccard distance based on accessory-gene presence/absence. The 27 strains together were compared as a broad-sense natto group. Horizontal gene transfer (HGT) and selective retention are plausible explanations, not mechanisms directly demonstrated here. Functional interpretations were based mainly on annotation and enrichment analyses, not direct phenotyping of multiple bekang isolates; the bekang comparison is an n=1 case study.
Comparative-genomics Workflow
The comparative-genomics workflow integrates public-data acquisition, genome assembly, quality control, and annotation with pangenome analysis and comparisons of phylogeny and gene content.
- Data preparation: Public raw reads and genome records were processed through consistent assembly, quality control, and annotation.
- Comparison of phylogeny and gene content: Core-genome phylogeny was combined with Jaccard-distance clustering based on accessory-gene presence/absence.
- Candidate-difference analysis: Synteny, phylogenetic networks, and group-specific genes were examined to compare candidate differences such as indels, transposons, gene duplications, and truncations.
Scope of Evidence and Potential Applications
The evidence identifies genomic relationships and candidate genes. Applying it to production performance, flavor, safety, or quality control requires phenotyping, fermentation experiments, reproducibility testing, and validation in manufacturing settings.
- Potential use: The workflow can support strain comparisons and generate hypotheses about genes or gene sets that may be associated with phenotypes.
- Validation required: It does not directly diagnose strain performance, product quality, or contamination from sequence data alone; experimental and manufacturing-level validation is required.
We welcome collaborations that combine food-microorganism genomes with phenotypic and fermentation data to test the valid scope of this comparative-genomics workflow.
Summary
By combining core-genome phylogeny with accessory-gene composition, the study presents a high-throughput comparative-genomics workflow that detects similarities among phylogenetically distinct B. subtilis strains. The bekang result is an n=1 case study, and its functional hypotheses require further validation.