プレプリント / バージョン1

Quantifying Cross-Species Similarity in Distributions of Amino Acid Compositional Divergence: Codon Partitioning and Pairwise Wasserstein Analysis of 85 Species

##article.authors##

DOI:

https://doi.org/10.51094/jxiv.5722

キーワード:

Amino acid composition、 Codon composition、 Synonymous codon usage、 Kullback–Leibler divergence、 Wasserstein distance

抄録

Our previous study showed that distributions of Jensen–Shannon distances between individual protein amino acid compositions and their species-specific mean compositions were unexpectedly similar across 85 diverse species, whereas analogous distributions based on CDS nucleotide composition were substantially more heterogeneous. Here, we examined how this contrast was expressed across nucleotide, codon, and amino acid representations derived from the same coding sequences, using a forward Kullback–Leibler (KL) framework that permits exact partitioning of total codon divergence into amino acid-composition and synonymous-codon-usage components.

We reanalyzed 2,095,689 retained CDS records from the same 85 species under an operational standard-code framework. Each CDS was represented by the frequencies of 61 standard sense codons, from which nucleotide and amino acid compositions were derived. Wasserstein-1 distances between species-specific empirical KL-divergence distributions were calculated for all 3,570 unique species pairs, first using the complete retained CDS set and then within three coding-length bins (100–199, 200–399, and 400–799 included codons). Pairwise distances were standardized by the median within-species interquartile range for the corresponding divergence type and analysis set.

In the whole-dataset analysis, the standardized mean pairwise distance was lowest for amino acid divergence (0.254), followed by total codon divergence (0.335), synonymous divergence (0.396), and nucleotide divergence (0.523). Amino acid divergence also showed the lowest standardized distance in every coding-length bin (0.258–0.311) and remained relatively stable across bins, whereas the standardized distances for the other divergence types increased with coding length.

Thus, the nucleotide–amino acid contrast previously observed using Jensen–Shannon distance was recapitulated under forward KL divergence and direct distribution-to-distribution comparison. Under the IQR-standardized comparison used here, cross-species similarity was greatest at the amino acid-composition level, although the relative contributions of codon-to-amino-acid dimensional reduction and biological constraints remain to be determined.

利益相反に関する開示

The author declare no conflicts of interest associated with this manuscript.

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

Sueoka, N. (1961). Compositional Correlation between Deoxyribonucleic Acid and Protein. Cold Spring Harbor Symposia on Quantitative Biology, 26(0), 35–43. https://doi.org/10.1101/SQB.1961.026.01.009

Knight, R. D., Freeland, S. J., & Landweber, L. F. (2001). A simple model based on mutation and selection explains trends in codon and amino-acid usage and GC composition within and across genomes. Genome Biology, 2(4), research0010.1. https://doi.org/10.1186/gb-2001-2-4-research0010

Singer, G. A. C., & Hickey, D. A. (2000). Nucleotide Bias Causes a Genomewide Bias in the Amino Acid Composition of Proteins. Molecular Biology and Evolution, 17(11), 1581–1588. https://doi.org/10.1093/oxfordjournals.molbev.a026257

Esumi, G. (2026). Proteome amino acid compositional variation is uniform across 85 species from the three domains of life [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.2928

Kullback, S., & Leibler, R. A. (1951). On Information and Sufficiency. The Annals of Mathematical Statistics, 22(1), 79–86. https://doi.org/10.1214/aoms/1177729694

Banerjee A, Merugu S, Dhillon IS, Ghosh J (2005) Clustering with Bregman divergences. Journal of Machine Learning Research 6:1705–1749. https://www.jmlr.org/papers/v6/banerjee05b.html

Peyré, G., & Cuturi, M. (2019). Computational Optimal Transport with Applications to Data Sciences. Foundations and Trends® in Machine Learning, 11(5–6), 355–607. https://doi.org/10.1561/2200000073

EMBL-EBI. (2024). Reference Proteomes (Release 2024_02) [Database]. Retrieved January 20, 2026, from https://www.ebi.ac.uk/reference_proteomes/

O’Leary, N. A., Wright, M. W., Brister, J. R., Ciufo, S., Haddad, D., McVeigh, R., Rajput, B., Robbertse, B., Smith-White, B., Ako-Adjei, D., Astashyn, A., Badretdin, A., Bao, Y., Blinkova, O., Brover, V., Chetvernin, V., Choi, J., Cox, E., Ermolaeva, O., … Pruitt, K. D. (2016). Reference sequence (RefSeq) database at NCBI: current status, taxonomic expansion, and functional annotation. Nucleic Acids Research, 44(D1), D733–D745. https://doi.org/10.1093/nar/gkv1189

Esumi, G. (2023). The Synonymous Codon Usage of a Protein Gene Is Primarily Determined by the Guanine + Cytosine Content of the Individual Gene Rather Than the Species to Which It Belongs To Synthesize Proteins With a Balanced Amino Acid Composition [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.561

Esumi, G. (2025). Amino‑acid composition is a selection target for coding‑sequence retention: evidence from out‑of‑frame translation comparisons in Escherichia coli [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.1565

Esumi, G. (2023). Statistical Extremes of Amino Acid Residue Composition of the Proteome Proteins Can Explain the Origin of the Universality of the Genetic Code [Preprint]. Jxiv. https://doi.org/10.51094/jxiv.575

ダウンロード

公開済


投稿日時: 2026-07-24 04:24:47 UTC

公開日時: 2026-08-03 07:33:40 UTC
研究分野
生物学・生命科学・基礎医学