Analyze Diet
Journal of equine veterinary science2026; 105924; doi: 10.1016/j.jevs.2026.105924

Accuracy of large language model-based artificial intelligence tools for equine topics.

Abstract: Artificial intelligence (AI) platforms are becoming increasingly popular as resources for equine information. However, these platforms generate responses from a wide range of sources and do not always distinguish between fact and opinion. Objective: The objective of this study was to assess the accuracy and quality of AI-generated answers to equine-related questions. Researchers hypothesized that AI platforms could answer basic equine questions effectively but would perform poorly on complex topics or questions. Methods: Forty questions were written covering general horse care, facilities management, nutrition, genetics, and reproduction. Each question was categorized by difficulty level: beginner, intermediate, advanced, or trending. Three AI platforms were tested: ChatGPT (CGPT), Microsoft Copilot (MicCP), and ExtensionBot (ExtBot). Responses were scored for accuracy, relevance, thoroughness, and source quality (5 points each; total 20). Data were analyzed using PROC GLM in SAS (v. 9.4). Results: Total score was affected by level (P = 0.002). Intermediate questions had the highest total score (15.95 ± 1.99). Accuracy was affected by platform (P < 0.001), level (P < 0.001), and topic (P = 0.015). CGPT (4.18 ± 0.93) and MicCP (4.08 ± 0.83) outperformed ExtBot (3.26 ± 1.21). Relevance was affected by platform (P = 0.042) and level (P < 0.001). Thoroughness was affected by platform (P < 0.001). Source quality differed by platform (P = 0.037). Conclusions: AI platforms could be resources; currently they fall short of the knowledge that Equine Extension Specialists can offer. AI platforms had difficulty addressing complex topics and demonstrated inconsistent performance across criteria.
Publication Date: 2026-05-01 PubMed ID: 42069023DOI: 10.1016/j.jevs.2026.105924Google Scholar: Lookup
The Equine Research Bank provides access to a large database of publicly available scientific literature. Inclusion in the Research Bank does not imply endorsement of study methods or findings by Mad Barn.
  • Journal Article

Summary

This research summary has been generated with artificial intelligence and may contain errors and omissions. Refer to the original study to confirm details provided. Submit correction.

Accuracy of large language model-based artificial intelligence tools for equine topics evaluated their performance in answering questions on horse care and management.

Objective and Hypothesis

  • The study aimed to assess the accuracy, relevance, thoroughness, and source quality of AI-generated responses on equine-related questions.
  • Researchers hypothesized that AI platforms would perform well on basic questions but poorly on complex topics.

Methods

  • Forty questions were designed covering diverse equine topics: general horse care, facilities management, nutrition, genetics, and reproduction.
  • Questions were classified by difficulty into four categories: beginner, intermediate, advanced, and trending.
  • Three AI platforms were tested:
    • ChatGPT (CGPT)
    • Microsoft Copilot (MicCP)
    • ExtensionBot (ExtBot)
  • Each AI’s response was scored on four criteria: accuracy, relevance, thoroughness, and source quality, with a maximum of 5 points each (total score 20).
  • Statistical analysis was performed using PROC GLM in SAS version 9.4 to determine differences by platform, question difficulty, and topic.

Key Results

  • Total Score and Difficulty:
    • Question difficulty affected the total score significantly (P = 0.002).
    • Intermediate-level questions received the highest total scores (mean 15.95 out of 20, SD 1.99).
  • Accuracy:
    • Platform, difficulty level, and topic significantly influenced accuracy (P < 0.001 for platform and level; P = 0.015 for topic).
    • ChatGPT scored highest on accuracy (4.18 ± 0.93), followed closely by Microsoft Copilot (4.08 ± 0.83).
    • ExtensionBot lagged behind with a lower accuracy score (3.26 ± 1.21).
  • Relevance:
    • Platform (P = 0.042) and question difficulty (P < 0.001) significantly affected relevance scores.
  • Thoroughness:
    • Platform had a significant effect on thoroughness (P < 0.001), indicating varying depth of responses.
  • Source Quality:
    • Significantly differed by platform (P = 0.037), highlighting variability in the reliability of referenced sources.

Conclusions

  • AI platforms demonstrate potential as informational resources for equine topics but currently do not match the expertise offered by Equine Extension Specialists.
  • Performance was inconsistent, particularly on complex or advanced topics where the AI struggled to provide accurate, thorough, and high-quality sources.
  • Among the tested platforms, ChatGPT and Microsoft Copilot generally outperformed ExtensionBot.
  • The findings suggest caution in relying solely on AI-generated information for complex equine care decisions and highlight the continued importance of expert human guidance.

Cite This Article

APA
Aldworth-Yang S, Coleman SJ, O'Reilly K, Catalano DN. (2026). Accuracy of large language model-based artificial intelligence tools for equine topics. J Equine Vet Sci, 105924. https://doi.org/10.1016/j.jevs.2026.105924

Publication

ISSN: 0737-0806
NlmUniqueID: 8216840
Country: United States
Language: English
Pages: 105924
PII: S0737-0806(26)00159-0

Researcher Affiliations

Aldworth-Yang, S
  • Department of Animal Sciences, College of Agricultural Sciences, Colorado State University, 350 W Pitkin Street, Fort Collins, Colorado, USA, 80521.
Coleman, S J
  • Department of Animal Sciences, College of Agricultural Sciences, Colorado State University, 350 W Pitkin Street, Fort Collins, Colorado, USA, 80521.
O'Reilly, K
  • Department of Animal Sciences, College of Agricultural Sciences, Colorado State University, 350 W Pitkin Street, Fort Collins, Colorado, USA, 80521.
Catalano, D N
  • Department of Animal Sciences, College of Agricultural Sciences, Colorado State University, 350 W Pitkin Street, Fort Collins, Colorado, USA, 80521. Electronic address: Devan.Catalano@colostate.edu.

Conflict of Interest Statement

Declaration of competing interest None of the authors has any financial or personal relationships that could inappropriately influence or bias the content of the paper.

Citations

This article has been cited 0 times.