Quantifying Cross-Species Similarity in Distributions of Amino Acid Compositional Divergence: Codon Partitioning and Pairwise Wasserstein Analysis of 85 Species
DOI:
https://doi.org/10.51094/jxiv.5722キーワード:
Amino acid composition、 Codon composition、 Synonymous codon usage、 Kullback–Leibler divergence、 Wasserstein distance抄録
Our previous study showed that distributions of Jensen–Shannon distances between individual protein amino acid compositions and their species-specific mean compositions were unexpectedly similar across 85 diverse species, whereas analogous distributions based on CDS nucleotide composition were substantially more heterogeneous. Here, we examined how this contrast was expressed across nucleotide, codon, and amino acid representations derived from the same coding sequences, using a forward Kullback–Leibler (KL) framework that permits exact partitioning of total codon divergence into amino acid-composition and synonymous-codon-usage components.
We reanalyzed 2,095,689 retained CDS records from the same 85 species under an operational standard-code framework. Each CDS was represented by the frequencies of 61 standard sense codons, from which nucleotide and amino acid compositions were derived. Wasserstein-1 distances between species-specific empirical KL-divergence distributions were calculated for all 3,570 unique species pairs, first using the complete retained CDS set and then within three coding-length bins (100–199, 200–399, and 400–799 included codons). Pairwise distances were standardized by the median within-species interquartile range for the corresponding divergence type and analysis set.
In the whole-dataset analysis, the standardized mean pairwise distance was lowest for amino acid divergence (0.254), followed by total codon divergence (0.335), synonymous divergence (0.396), and nucleotide divergence (0.523). Amino acid divergence also showed the lowest standardized distance in every coding-length bin (0.258–0.311) and remained relatively stable across bins, whereas the standardized distances for the other divergence types increased with coding length.
Thus, the nucleotide–amino acid contrast previously observed using Jensen–Shannon distance was recapitulated under forward KL divergence and direct distribution-to-distribution comparison. Under the IQR-standardized comparison used here, cross-species similarity was greatest at the amino acid-composition level, although the relative contributions of codon-to-amino-acid dimensional reduction and biological constraints remain to be determined.
利益相反に関する開示
The author declare no conflicts of interest associated with this manuscript.ダウンロード *前日までの集計結果を表示します
引用文献
Sueoka, N. (1961). Compositional Correlation between Deoxyribonucleic Acid and Protein. Cold Spring Harbor Symposia on Quantitative Biology, 26(0), 35–43. https://doi.org/10.1101/SQB.1961.026.01.009
Knight, R. D., Freeland, S. J., & Landweber, L. F. (2001). A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes. Genome Biology, 2(4), research0010.1. https://doi.org/10.1186/gb-2001-2-4-research0010
Singer, G. A. C., & Hickey, D. A. (2000). Nucleotide Bias Causes a Genomewide Bias in the Amino Acid Composition of Proteins. Molecular Biology and Evolution, 17(11), 1581–1588. https://doi.org/10.1093/oxfordjournals.molbev.a026257
Esumi, G. (2026). Proteome amino acid compositional variation is uniform across 85 species from the three domains of life [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.2928
Kullback, S., & Leibler, R. A. (1951). On Information and Sufficiency. The Annals of Mathematical Statistics, 22(1), 79–86. https://doi.org/10.1214/aoms/1177729694
Banerjee A, Merugu S, Dhillon IS, Ghosh J (2005) Clustering with Bregman divergences. Journal of Machine Learning Research 6:1705–1749. https://www.jmlr.org/papers/v6/banerjee05b.html
Peyré, G., & Cuturi, M. (2019). Computational Optimal Transport with Applications to Data Sciences. Foundations and Trends® in Machine Learning, 11(5–6), 355–607. https://doi.org/10.1561/2200000073
EMBL-EBI. (2024). Reference Proteomes (Release 2024_02) [Database]. Retrieved January 20, 2026, from https://www.ebi.ac.uk/reference_proteomes/
O’Leary, N. A., Wright, M. W., Brister, J. R., Ciufo, S., Haddad, D., McVeigh, R., Rajput, B., Robbertse, B., Smith-White, B., Ako-Adjei, D., Astashyn, A., Badretdin, A., Bao, Y., Blinkova, O., Brover, V., Chetvernin, V., Choi, J., Cox, E., Ermolaeva, O., … Pruitt, K. D. (2016). Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research, 44(D1), D733–D745. https://doi.org/10.1093/nar/gkv1189
Esumi, G. (2023). The Synonymous Codon Usage of a Protein Gene Is Primarily Determined by the Guanine + Cytosine Content of the Individual Gene Rather Than the Species to Which It Belongs To Synthesize Proteins With a Balanced Amino Acid Composition [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.561
Esumi, G. (2025). Amino‑acid composition is a selection target for coding‑sequence retention: evidence from out‑of‑frame translation comparisons in Escherichia coli [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.1565
Esumi, G. (2023). Statistical Extremes of Amino Acid Residue Composition of the Proteome Proteins Can Explain the Origin of the Universality of the Genetic Code [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.575
ダウンロード
公開済
投稿日時: 2026-07-24 04:24:47 UTC
公開日時: 2026-08-03 07:33:40 UTC
ライセンス
Copyright(c)2026
Genshiro Esumi
この作品は、Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licenseの下でライセンスされています。
