Analyze Diet
BMC genomics2026; doi: 10.1186/s12864-026-12862-0

Genomic diversity and structure in arabian horses revealed by whole-genome sequencing: establishment of an allele frequency database of common genetic variation.

Abstract: BACKGROUND: The Arabian horse is a culturally and historically influential breed that has contributed to the development of many modern horse populations. However, genomic resources for this breed remain limited, particularly population-level allele frequency datasets that support studies of genetic diversity, selection, and disease. The aim of this study was to generate a comprehensive allele frequency database for Arabian horses using representative sampling and high-coverage whole-genome sequencing. RESULTS: We generated the first publicly available allele frequency database for the Arabian horse by sequencing pooled genomic DNA from 120 individuals originating from four distinct populations: Poland, the United States, Egypt, and Syria. Across the dataset, 11.7 million single nucleotide polymorphisms were identified, and both population-specific and global allele frequencies were estimated. Genome-wide analyses revealed substantial variation in polymorphism density, as well as multiple regions exhibiting fixation for alternate alleles. Principal component analysis, hierarchical clustering, and admixture analysis demonstrated clear genetic differentiation among the four populations. We detected 6,210 fixed genomic blocks, many of which overlapped with genes involved in metabolic and signaling pathways, suggesting potential historical selection. One such block corresponded to a region previously associated with insect bite hypersensitivity. To assess clinical relevance, we screened the dataset for known pathogenic variants linked to four inherited disorders described in Arabian horses. None of these variants were detected, and their absence was supported by high sequencing depth and confirmed through Sanger sequencing. CONCLUSIONS: This study provides the first comprehensive allele frequency resource for the Arabian horse. The database enables robust variant filtering, facilitates the identification of candidate functional polymorphisms, and supports studies of genetic structure and selection. It represents an important foundation for future work in equine genomics and has practical applications in breeding, conservation genetics, and health management. Continued expansion of the dataset with additional populations will further enhance its value for research and applied genomics.
Publication Date: 2026-04-20 PubMed ID: 42010471DOI: 10.1186/s12864-026-12862-0Google Scholar: Lookup
The Equine Research Bank provides access to a large database of publicly available scientific literature. Inclusion in the Research Bank does not imply endorsement of study methods or findings by Mad Barn.
  • Journal Article

Summary

This research summary has been generated with artificial intelligence and may contain errors and omissions. Refer to the original study to confirm details provided. Submit correction.

This study sequenced pooled DNA from 120 Arabian horses across four countries to catalog 11.7 million common genetic variants and build the first breed-wide allele frequency database. The data reveal clear genetic structure among populations, regions of the genome likely shaped by selection, and provide a practical resource for breeding, conservation, and health research.

Background and rationale

  • Arabian horses have had a major historical and genetic influence on many modern horse breeds, yet breed-specific genomic resources have lagged behind those for other equines.
  • Allele frequency databases are essential for distinguishing common from rare variants, prioritizing candidate variants in disease and trait studies, and understanding population structure and selection.
  • This work set out to create a comprehensive, publicly available allele frequency resource for Arabian horses using representative global sampling and high-coverage whole-genome sequencing.

Study design and methods

  • Sampling strategy: Pooled genomic DNA from 120 Arabian horses spanning four distinct populations (Poland, United States, Egypt, Syria) to capture geographically and historically important lineages.
  • Sequencing and variant discovery: High-coverage whole-genome sequencing of pools; identification of 11.7 million single nucleotide polymorphisms (SNPs) across the genome.
  • Allele frequency estimation: Computed both population-specific and global allele frequencies from pooled read counts, enabling comparisons within and across populations.
  • Population structure analyses: Applied principal component analysis (PCA), hierarchical clustering, and admixture analysis to quantify genetic differentiation and shared ancestry.
  • Genomic landscape characterization: Assessed genome-wide polymorphism density and identified regions showing fixation for alternate alleles (i.e., all sampled chromosomes carrying the same non-reference allele).
  • Fixed-block mapping and functional context: Detected 6,210 fixed genomic blocks and overlapped these with annotated genes and pathways to infer potential biological relevance.
  • Clinical screening: Queried the dataset for previously reported pathogenic variants underlying four inherited disorders described in Arabian horses; validated absence by leveraging deep coverage and Sanger sequencing.

Main findings

  • Variant catalog: 11.7 million SNPs were discovered, providing a dense and breed-relevant map of common genetic variation.
  • Heterogeneity in diversity: Polymorphism density varied across the genome, highlighting regions of constraint or historical selection versus regions tolerant of variation.
  • Fixed regions: Multiple genomic segments showed fixation for alternate alleles; in total, 6,210 fixed blocks were identified.
  • Functional overlap: Many fixed blocks intersected genes involved in metabolic and signaling pathways, consistent with selection on performance, physiology, or adaptation.
  • Trait-relevant signal: One fixed block corresponded to a region previously linked to insect bite hypersensitivity, supporting biological plausibility of selection signals.
  • Population structure: PCA, clustering, and admixture results showed clear genetic differentiation among Polish, U.S., Egyptian, and Syrian populations, alongside measurable shared ancestry patterns.
  • Clinical variant screen: None of the known pathogenic variants for four Arabian inherited disorders were detected; high sequencing depth and Sanger confirmation substantiate their absence in the sampled cohorts.
  • Resource delivery: The study provides the first publicly available, population-level allele frequency database specific to the Arabian breed.

Interpretation and implications

  • Selection and history: Concentrated fixed blocks overlapping key pathways suggest historical selection, whether for breed characteristics, environmental adaptation, or management practices.
  • Population differentiation: The genetic distinctiveness among national populations reflects founder effects, breeding policies, and limited gene flow; admixture patterns document historical exchanges.
  • Health genetics: A reliable allele frequency background enables robust filtering of candidate variants in veterinary diagnostics and research, reducing false positives from common benign variants.
  • Conservation and breeding: Frequency information helps identify rare alleles, monitor inbreeding, preserve diversity across subpopulations, and avoid inadvertent fixation of deleterious alleles.
  • Research acceleration: The database supports genome scans for selection, genotype–phenotype mapping, and comparative studies across horse breeds.

Technical notes and definitions

  • Allele frequency: The proportion of chromosomes carrying a given variant in a population; critical for distinguishing common background variation from potentially causal, rare changes.
  • Fixation: A variant is “fixed” when essentially all sampled chromosomes carry the same allele; blocks of adjacent fixed SNPs can indicate selective sweeps or strong drift.
  • PCA/admixture: PCA summarizes genetic variation along major axes to show clustering; admixture estimates the proportion of ancestry from multiple genetic sources.
  • Pooled sequencing: Sequencing mixed DNA reduces costs and yields accurate population allele frequencies for common variants but does not provide per-individual genotypes or haplotypes.
  • Depth-supported absence: High read depth across sites increases confidence that clinically important variants are truly absent rather than missed due to insufficient coverage; Sanger validation adds orthogonal confirmation.

Strengths

  • First comprehensive, breed-specific allele frequency resource for Arabian horses, spanning multiple geographically distinct populations.
  • High-coverage sequencing and cross-validation (including Sanger) enhance data quality and clinical reliability.
  • Integration of population genetics and functional annotation contextualizes signals of selection and trait relevance.

Limitations

  • Pooled design limits resolution for individual genotypes, phasing, and very rare variants; allele frequency estimates are most accurate for common variants.
  • Sampling covers four populations; additional regions/lineages may harbor distinct variation not captured here.
  • Focus on SNPs likely underrepresents structural variants, copy-number changes, and indels, which also contribute to phenotype.
  • Signals of selection inferred from fixation and differentiation require cautious interpretation and benefit from complementary evidence (e.g., extended haplotype tests, functional assays).

Future directions

  • Expand sampling to additional Arabian subpopulations and bloodlines (e.g., other Middle Eastern, European, and Latin American lineages) to refine global allele frequencies.
  • Incorporate individual-level sequencing to capture rare variants, enable haplotype-based selection scans, and improve demographic inference.
  • Augment with non-SNP variation (indels, structural variants) and transcriptomic/epigenomic data for functional interpretation.
  • Link genotype data with detailed phenotypes (performance, conformation, disease) to discover genotype–phenotype associations.
  • Establish routine updates and community deposition pipelines to keep the database current and maximally representative.

Practical applications

  • Breeders: Use allele frequency data to maintain genetic diversity, minimize inbreeding, and avoid concentrating undesirable alleles.
  • Veterinarians and diagnostic labs: Filter variants in clinical sequencing against breed-specific frequencies to prioritize likely pathogenic candidates.
  • Conservation programs: Monitor population structure and diversity to guide mating plans and preserve distinct genetic lineages.
  • Researchers: Perform selection scans, admixture mapping, and cross-breed comparisons using a validated Arabian reference frequency panel.

Bottom line

  • This study delivers a foundational allele frequency database for Arabian horses, reveals meaningful population structure and selection-linked genomic regions, and provides an immediately useful toolset for breeding, conservation, and health genomics. Ongoing expansion across additional populations will further increase its accuracy and utility.

Cite This Article

APA
Szmatoła T, Finno C, Gurgul A, Heath H, Stefaniuk-Szmukier M, Almarzook S, Norton E, Ropka-Molik K. (2026). Genomic diversity and structure in arabian horses revealed by whole-genome sequencing: establishment of an allele frequency database of common genetic variation. BMC Genomics. https://doi.org/10.1186/s12864-026-12862-0

Publication

ISSN: 1471-2164
NlmUniqueID: 100965258
Country: England
Language: English

Researcher Affiliations

Szmatoła, Tomasz
  • Department of Basic Sciences, University of Agriculture in Kraków, Kraków, Poland.
  • Department of Animal Molecular Biology, Laboratory of Genomics, National Research Institute of Animal Production, Krakowska 1, Kraków, 32-083 Balice, Poland.
Finno, Carrie
  • Department of Population Health and Reproduction, Davis School of Veterinary Medicine, University of California, Davis, USA.
Gurgul, Artur
  • Department of Basic Sciences, University of Agriculture in Kraków, Kraków, Poland.
Heath, Harrison
  • Department of Biomolecular Engineering and Bioinformatics, University of California, Santa Cruz, USA.
Stefaniuk-Szmukier, Monika
  • Department of Animal Molecular Biology, Laboratory of Genomics, National Research Institute of Animal Production, Krakowska 1, Kraków, 32-083 Balice, Poland.
Almarzook, Saria
  • Department of Health and Care Management, Arden University, Berlin Campus, Berlin, Germany.
Norton, Elaine
  • College of Veterinary Medicine, University of Arizona, Oro Valley, USA.
Ropka-Molik, Katarzyna
  • Department of Animal Molecular Biology, Laboratory of Genomics, National Research Institute of Animal Production, Krakowska 1, Kraków, 32-083 Balice, Poland. katarzyna.ropka@iz.edu.pl.

Grant Funding

  • 503-182-609 / the Ministry of Agriculture and Rural Development of the Republic of Poland

Conflict of Interest Statement

Declarations. Ethics approval and consent to participate: The material used in this study comprised hair follicles and blood samples that originated from the genetic material bank of the National Research Institute of Animal Production (Balice, Poland). The samples were obtained and stored in accordance with the applicable Polish law, including the Act of 15 January 2015 on the Protection of Animals Used for Scientific or Educational Purposes (Journal of Laws 2015, item 266). The samples used as archival material from the Biological Material Bank at the National Research Institute of Animal Production were collected with the approval of the Institutional Animal Care and Use Committee (no. 1173/2015 and 102/2025). Hair follicle and blood samples were collected with informed owner consent. Competing interests: The authors declare no competing interests.

References

This article includes 35 references
  1. Cosgrove EJ, Sadeghi R, Schlamp F, Holl HM, Moradi-Shahrbabak M, Miraei-Ashtiani SR. Genome diversity and the origin of the Arabian horse. Sci Rep 2020;10:9702.
  2. Petersen JL, Mickelson JR, Cothran EG, Andersson LS, Axelsson J, Bailey E. Genetic diversity in the modern horse illustrated from genome-wide SNP data. PLoS ONE 2013;8(1):e54997.
  3. Librado P, Fages A, Gaunitz C, Khan N, Schubert M, Albrechtsen A. Ancient genomic changes associated with domestication of the horse. Science 2017;356(6336):442–5.
    doi: 10.1126/science.aam5298google scholar: lookup
  4. Orlando L, Ginolhac A, Zhang G, Froese D, Albrechtsen A, Stiller M. Recalibrating Equus evolution using the genome sequence of an early Middle Pleistocene horse. Nature 2013;499(7456):74–8.
    doi: 10.1038/nature12323google scholar: lookup
  5. Karczewski KJ, Francioli LC, Tiao G, Cummings BB, Alföldi J, Wang Q. The mutational constraint spectrum quantified from variation in 141,456 humans. Nature 2020;581:434–43.
    doi: 10.1038/s41586-020-2308-7google scholar: lookup
  6. Lek M, Karczewski KJ, Minikel EV, Samocha KE, Banks E, Fennell T. Analysis of protein-coding genetic variation in 60,706 humans. Nature 2016;536:285–91.
    doi: 10.1038/nature19057google scholar: lookup
  7. Brooks SA, Gabreski NA, Miller DC, Brisbin A, Brown HE, Streeter C. Whole-genome SNP association in the horse: identification of a deletion in myosin Va responsible for lavender foal syndrome. PLoS Genet 2010;6(4):e1000909.
  8. Aleman M, Finno CJ, Weich K, Penedo MCT. Investigation of known genetic mutations of Arabian horses in Egyptian Arabian foals with juvenile idiopathic epilepsy. J Vet Intern Med 2017;32(1):465–8.
    doi: 10.1111/jvim.14873google scholar: lookup
  9. Khanshour A, Conant E, Juras R, Cothran EG. Microsatellite analysis of genetic diversity and population structure of Arabian horse populations. J Hered 2013;104(3):386–98.
    doi: 10.1093/jhered/est003google scholar: lookup
  10. Klecel W, Głażewska I. Population structure and genetic diversity of Polish Arabian horses based on pedigree data. Animal 2023;17(3):100679.
  11. Wallner B, Vogl C, Shukla P, Burgstaller JP, Druml T, Brem G. Y chromosome uncovers the recent oriental origin of modern stallions. Curr Biol 2017;27(13):2029–e355.
    doi: 10.1016/j.cub.2017.05.086google scholar: lookup
  12. Andrews S. FastQC: a quality control tool for high throughput sequence data. 2010.
  13. Dodt M, Roehr JT, Ahmed R, Dieterich C. FLEXBAR—Flexible Barcode and Adapter Processing for Next-Generation Sequencing Platforms. Biology (Basel) 2012;1(3):895–905.
    doi: 10.3390/biology1030895google scholar: lookup
  14. Li H. Aligning sequence reads, clone sequences and assembly contigs with BWA-MEM. arXiv [q-bio.GN] 2013.
  15. Li H, Handsaker B, Wysoker A, Fennell T, Ruan J, Homer N. The Sequence Alignment/Map format and SAMtools. Bioinformatics 2009;25(16):2078–9.
  16. McKenna A, Hanna M, Banks E, Sivachenko A, Cibulskis K, Kernytsky A. The Genome Analysis Toolkit: a MapReduce framework for analyzing next-generation DNA sequencing data. Genome Res 2010;20(9):1297–303.
    doi: 10.1101/gr.107524.110google scholar: lookup
  17. Smit AFA, Hubley R, Green P. RepeatMasker Open-4.0. 2013–2015.
  18. R Core Team. R: A language and environment for statistical computing. Vienna: R Foundation for Statistical Computing. 2023.
  19. Paradis E, Schliep K. ape 5.0: an environment for modern phylogenetics and evolutionary analyses in R. Bioinformatics 2019;35(3):526–8.
  20. Alexander DH, Novembre J, Lange K. Fast model-based estimation of ancestry in unrelated individuals. Genome Res 2009;19(9):1655–64.
    doi: 10.1101/gr.094052.109google scholar: lookup
  21. Bu D, Luo H, Huo P, Wang Z, Zhang S, He Z. KOBAS-i: intelligent prioritization and exploratory visualization of biological functions for gene enrichment analysis. Nucleic Acids Res 2021;49(W1):W317–25.
    doi: 10.1093/nar/gkab447google scholar: lookup
  22. Quinlan AR, Hall IM. BEDTools: a flexible suite of utilities for comparing genomic features. Bioinformatics 2010;26(6):841–2.
  23. Ayad A, Almarzook S, Besseboua O, Aissanou S, Piórkowska K, Musiał AD. Investigation of Cerebellar Abiotrophy (CA), Lavender Foal Syndrome (LFS), and Severe Combined Immunodeficiency (SCID) variants in a cohort of three MENA region horse breeds. Genes (Basel) 2021;12(12):1893.
    doi: 10.3390/genes12121893google scholar: lookup
  24. Finno CJ, Spier SJ, Valberg SJ. Equine diseases caused by known genetic mutations. Vet J 2009;179(3):336–47.
  25. Scott E, Woolard K, Finno CJ, Murray JD. Cerebellar abiotrophy across domestic species. Cerebellum 2018;17(3):372–9.
    doi: 10.1007/s12311-017-0914-1google scholar: lookup
  26. Finno CJ. Occipitoatlantoaxial malformations in equids. Equine Vet Educ 2025.
    doi: 10.1111/eve.14160google scholar: lookup
  27. Doan R, Cohen ND, Sawyer J, Ghaffari N, Johnson CD, Dindot SV. Identification of copy number variants in horses. Genome Res 2012;22(5):899–907.
    doi: 10.1101/gr.128991.111google scholar: lookup
  28. Daetwyler HD, Capitan A, Pausch H, Stothard P, van Binsbergen R, Brøndum RF. Whole-genome sequencing of 234 bulls facilitates mapping of monogenic and complex traits in cattle. Nat Genet 2014;46(8):858–65.
    doi: 10.1038/ng.3034google scholar: lookup
  29. Bers DM. Cardiac excitation–contraction coupling. Nature 2002;415(6868):198–205.
    doi: 10.1038/415198agoogle scholar: lookup
  30. Schröder W, Klostermann A, Distl O. Candidate genes for physical performance in the horse. Vet J 2011;190(1):39–48.
  31. Signer-Hasler H, Flury C, Haase B, Burger D, Simianer H, Leeb T. A genome-wide association study reveals loci influencing height and other conformation traits in horses.. PLoS ONE 2012;7(5):e37282.
  32. Akiyama H, Chaboissier M, Martin JF, Schedl A, de Crombrugghe B. The transcription factor SOX9 has essential roles in successive steps of the chondrocyte differentiation pathway and is required for expression of SOX5 and SOX6.. Genes Dev 2002;16(21):2813–28.
    doi: 10.1101/gad.1017802google scholar: lookup
  33. Schurink A, Wolc A, Ducro B, Frankena K, Garrick DJ, Dekkers JCM. Genome-wide association study of insect bite hypersensitivity in two horse populations in the Netherlands.. Genet Sel Evol 2012;44(1):31.
    doi: 10.1186/1297-9686-44-31google scholar: lookup
  34. Grilz-Seger G, Neuditschko M, Ricard A, Velie B, Lindgren G, Mesarič M. Genome-wide homozygosity patterns and evidence for selection in a set of European and Near Eastern horse breeds.. Genes (Basel) 2019;10(7):491.
    doi: 10.3390/genes10070491google scholar: lookup
  35. Schlötterer C, Tobler R, Kofler R, Nolte V. Sequencing pools of individuals—mining genome-wide polymorphism data without big funding.. Nat Rev Genet 2014;15(11):749–63.
    doi: 10.1038/nrg3803google scholar: lookup

Citations

This article has been cited 0 times.