Abstract
Objective Efficient and high-quality history taking is central to vestibular diagnosis, but it is often constrained by limited consultation time and variable patient understanding of questionnaires. Large language models (LLMs) can help structure interviews and support diagnosis by proposing differential diagnoses; however, evidence of their use in neurotology remains scarce. Here, we developed and evaluated an LLM-based system that automatically conducts pre-consultation interviews for dizziness symptoms. Methods We developed a prototype interview application with text input and speech output. The patients spoke freely, and an examiner transcribed their utterances verbatim into the interface. The transcribed text was processed by an LLM via the OpenAI Application Programming Interface, which returned a textual reply. This reply was then converted into audio and played to the patient, enabling a natural voice-based exchange while maintaining text-only input to the model. The prompt was constructed based on a 14-item vertigo questionnaire routinely used in clinical practice, and its interview logic aligned with the diagnostic criteria of the Bárány Society and Japan Society for Equilibrium Research. Nineteen consecutive patients at a specialty clinic completed the interview prior to the otolaryngologist’s assessment. The primary outcomes were (1) completion of prespecified history items, (2) concordance between the LLM’s ranked differentials and the final chart diagnosis (Top-1 and Top-3), and (3) communication quality assessed using the Kalamazoo Essential Elements Communication Checklist–Adapted. We used Generative Pre-trained Transformer 4 (GPT-4) for the first seven interviews and GPT-4o for the subsequent twelve. Results The mean item completion rate was 88.7%. Items related to migraine were under-ascertained. The final diagnosis matched the system’s Top-1 suggestion in 9/19 cases (47.4%, 95% confidence interval [CI] 27.3–68.3) and was included within the Top-3 diagnoses in 15/19 cases (78.9%, 95% CI 56.7–91.5). Concordance was higher for Menière’s disease (Top-1 71.4%, Top-3 85.7%) and persistent postural-perceptual dizziness (PPPD; 66.7%, 100%) but lower for chronic unilateral vestibulopathy (20%, 40%), which generally requires physiological testing and lacks widely accepted diagnostic criteria. Communication quality ratings showed relatively higher scores for information gathering but lower scores for empathy and elicitation of the patient’s perspective. Conclusion This pilot study provides preliminary evidence for the feasibility of LLM-mediated patient interviews in vestibular medicine. The system showed promising diagnostic alignment for common conditions such as Ménière’s disease and PPPD. Studies with larger cohorts conducted in parallel with the continuing advances in LLMs are required to establish their clinical applicability.
| Original language | English (US) |
|---|---|
| Pages (from-to) | 423-428 |
| Number of pages | 6 |
| Journal | Auris Nasus Larynx |
| Volume | 53 |
| Issue number | 3 |
| DOIs | |
| State | Published - Jun 2026 |
| Externally published | Yes |
Keywords
- GPT-4
- Interview quality assessment
- Large language models
- Medical history taking
ASJC Scopus subject areas
- Surgery
- Otorhinolaryngology
Fingerprint
Dive into the research topics of 'Pilot evaluation of a large language model-mediated pre-consultation interview for vestibular disorders'. Together they form a unique fingerprint.Cite this
- APA
- Standard
- Harvard
- Vancouver
- Author
- BIBTEX
- RIS