Academic papers

Papers brief: PetQA tests whether AI can answer Korean vet questions like a clinic

arXiv PetQA benchmarks 18 LLMs and LVLMs on 18,827 Korean long-form dog-and-cat QA pairs from real pet-owner queries — text and multimodal — with expert veterinarian answers.

  • academic papers
  • Korean NLP
  • veterinary AI
  • LLM benchmarks

Source: arXiv

Paper

PetQA: Benchmarking Veterinary Knowledge and Clinical Reasoning — Taegyun Kim, Youngwook Ham, Jungwook Rhim, Ju-Hyun An, Sungkyu Park, Kunwoo Park (Soongsil University, Kangwon National University, KDI School of Public Policy and Management; arXiv 2609.04598, Sep 2026).

What it claims

Pet owners already ask chatbots about limps, vomiting, and odd behavior before they call a clinic. Most medical-AI benchmarks still target human medicine in English or Chinese multiple-choice formats — not the long-form, symptom-heavy questions people actually post about dogs and cats.

PetQA fills that gap with a Korean long-form QA benchmark built from real-world pet-health questions and answers verified by licensed veterinarians on a major Korean Q&A platform. The release contains 10,076 text-only pairs and 8,751 multimodal pairs — 18,827 total — focused on dogs and cats. The held-out PetQA-Bench test split adds labels for question type and clinical condition (2,000 items per modality in the test set).

The authors benchmark 18 models under zero-shot, retrieval-augmented generation (RAG), and supervised fine-tuning (SFT), scoring with ROUGE, BERTScore, and LLM-as-a-judge factuality and helpfulness. PetQA-Bench also ships in five translated languages via GitHub (ssu-humane/PetQA).

Headline findings: closed models generally lead on factuality and helpfulness; every model drops on multimodal items; RAG/SFT do not reliably lift judge scores.

The breakdown

PetQA is infrastructure, not a product launch. The design choices matter more than any single leaderboard rank.

Source realism: Items come from years of Korean pet-owner posts — not licensing-exam stems — matching how owners query AI before booking a visit.

Long-form over MCQA: Expert open answers stress reasoning, not pattern-matching among four choices.

Multimodal gap: Photos of rashes or posture often accompany real questions; the paper reports a consistent multimodal penalty.

Adaptation skepticism: RAG and SFT help sometimes on n-gram metrics but not consistently on factuality/helpfulness judges.

Why readers outside the lab should care

Korea pet-market teams: Naver-style community Q&A is where many owners first describe symptoms in Korean. If you ship a pet-health chatbot, tele-vet triage, or insurer symptom checker for the domestic market, PetQA is the first large-scale Korean long-form yardstick — not MedQA or KorMedMCQA repurposed for animals.

Overseas readers with pets in Korea: Expats often lean on English-first models trained on U.S. vet content. PetQA’s Korean expert references and five-language bench translations test whether your stack handles Korean clinical pet language, not just small talk.

Global vet-AI builders: The authors position PetQA as the first long-form veterinary QA benchmark — note the multimodal failure mode before marketing photo-based diagnosis.

What builders and Korea-touching teams should watch

  • Do evaluate on long-form Korean pet QA with expert references — not human-med MCQA or English-only pet blogs.
  • Don’t treat strong text scores as proof the model handles owner-uploaded photos; multimodal items lag across all eighteen models in the paper.
  • Expect RAG and SFT to be uneven on factuality/helpfulness judges; plan domain adaptation beyond generic retrieval or a shallow fine-tune.
  • Re-check any consumer pet-health bot marketed in Korea against real community question styles (symptom narrative, not trivia).
  • Use the released multilingual PetQA-Bench if you serve bilingual households — Korean source quality with translated test splits for cross-locale regression.

Context

Read PetQA as a Korean reality check for vet-facing AI, not proof that chatbots should replace clinics. Korelay’s frame: owners already ask the internet before the vet; this benchmark shows current models still stumble on the photos and long answers that real Korean pet posts require — and that bolt-on RAG or SFT is not a automatic fix.

Source

Primary: arXiv:2609.04598 (abstract and framing cited; open the OA PDF for methods and full results). Do not republish the PDF.