A study published on October 8 in The Lancet revealed that artificial intelligence chatbots have the potential to assist doctors in understanding patients’ conditions and preparing for consultations before appointments. However, researchers emphasized that current evidence is still insufficient to support AI independently diagnosing, triaging, or managing treatments.
The study was conducted in collaboration between Google researchers and the Beth Israel Deaconess Medical Center in Boston, USA. It involved 114 patients, with 100 engaging in consultations with the AI medical dialogue system AMIE and 98 subsequently seeing a physician.
Patients first interacted with AMIE through video calls, with the process supervised by operators in real-time. After completing the AI consultation, patients either visited a primary care physician in person or through telehealth within five days.
Throughout the study period, no AI dialogues were interrupted due to breaching predefined safety standards. Nevertheless, this does not mean that AI consultations are entirely risk-free.
Researchers reviewed medical records 8 weeks post-consultation, comparing the differential diagnoses proposed by AMIE with the final diagnoses. The results showed that the final diagnosis appeared as the top choice suggested by AMIE in 56.1% of cases; considering the top three potential diagnoses increased the percentage to 74.5%; and expanding to the top seven reached 89.8%.
In terms of doctors’ evaluations, in 44 cases where doctors had the opportunity to review the AMIE dialogue records or summaries beforehand, 75% acknowledged that AI helped in preparing for consultations. In 57% of cases, doctors believed that reviewing the AI consultation content in advance could potentially alter their subsequent actions or treatment decisions. The study also found that patients’ attitudes towards using AI in primary healthcare improved after interacting with AI.
However, a report from The Independent highlighted instances in the study where AI generated false information (hallucinations). In one conversation, AI listed cancer as a possible diagnosis, causing anxiety in the patient, with the doctor later rating the encounter as “somewhat harmful.” In another dialogue, the chatbot also produced incorrect information. The report also noted that in five interactions, clinical information was supplemented without specifying who provided it or the specific content.
Professor Victoria Tzortziou Brown, President of the Royal College of General Practitioners, expressed that the study’s scale was small and conducted only at a single healthcare institution in the US. Each AI conversation was closely monitored by doctors, and patients later engaged in discussions with physicians, making it challenging to conclude the safety and effectiveness of AI consultations without such supervision. She stressed that doctors’ clinical judgment, understanding of patients’ medical histories and life circumstances, and accumulated knowledge from long-term care remain crucial in diagnosis.
Researchers stated that further large-scale studies are necessary in the future to assess the safety of AI diagnostics, communication between patients and doctors, physician workload, costs, and the necessary supervision level to ensure safety.
AMIE (Articulate Medical Intelligence Explorer) is an AI medical dialogue system developed by Google’s research team, aiming to gather medical histories, analyze symptoms, and suggest potential diagnoses through conversations with patients for discussion during consultations with doctors.
Previous studies on AMIE primarily evaluated its performance in simulated consultation environments and compared it with primary care physicians. This study advanced to test the feasibility of patient-AI dialogues, dialogue safety, and whether the information organized by AI can aid doctors in preparing for consultations in real outpatient processes.
