Announced 8 Oct 2026 · Sources checked
What was published on 8 October?
The Lancet listing is dated 8 October 2026: “Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study.” The abstract we inspected describes a single-arm feasibility study at one academic clinic, registered as NCT06911398. Google’s research blog, signed by Mike Schaekermann and timestamped the same day, calls it “Google’s first-ever publication in the main journal of The Lancet.” BIDMC’s Newswise note, posted at 20:10 EDT on 8 October, is the hospital’s version of the same results.
A longer methods manuscript with the same title is on arXiv as 2603.08448. That HTML copy still says “100 adult patients completed” the chat. The Lancet abstract and BIDMC say 114 enrolled and 98 completed both the AMIE interaction and the appointment. We are dating the news to the 8 October journal and hospital notes, and using the preprint only where it explains methods that the abstract omits. The full Lancet HTML returned HTTP 403 to our fetch; the counts below are from the abstract, BIDMC, Google, the trial record, and the preprint.
This is not a UK-style deployment approval. For a regulator-facing healthcare process that is still in trial phases, see the UK AI healthcare commission’s Airlock Phase 3 note. AMIE here is a study intervention inside one Boston practice.
What did the 98 patients actually do?
ClinicalTrials.gov, last updated 20 April 2026 in the record we pulled, titles the protocol “AMIE's Clinical Conversational Abilities in an Urgent Care Setting.” Patients aged 18 or over who already had an episodic visit booked at Healthcare Associates first joined a secure chat. The registry says they shared their screen on a HIPAA-compliant video call while a physician safety supervisor watched. AMIE took a history, offered possible diagnoses for the patient to discuss with their doctor, and produced a transcript and summary for the primary-care clinician. The chat could happen up to five days before the visit, less than a week in the registry’s wording.
Inclusion required English as a primary language in the record. The registry excludes pregnancy-related visits, psychiatric concerns, visits to establish care, and follow-ups. Enrollment on the registry is listed as 100 actual. That is the planned-to-actual figure on ClinicalTrials.gov, not a contradiction we can resolve without the paywalled Lancet tables. BIDMC names Healthcare Associates as the clinic.
The preprint says the deployed AMIE build used Gemini 2.5 models with thinking mode and a “state-aware chain-of-reasoning” tuned for pre-visit history taking. Management plans were generated for research scoring and, in both the registry and the preprint, were not shown to patients or to the treating clinicians. We did not use the chatbot.
What do the safety and diagnosis numbers actually say?
The primary safety result is simple: zero conversation stops under the pre-specified rules. BIDMC lists the triggers as harm risk, significant distress, a patient asking to end, a need to clarify symptoms, emergency instructions, or concern about a serious diagnosis. Supervisors still recorded one hallucination and added clinical information in five chats. After the chat, the registry says, the supervisor was to talk with the patient and correct errors. “No safety stops” is therefore not “no human touched the transcript.”
On diagnosis, Google’s blog writes that “AMIE’s differential diagnoses also matched the doctors’ final diagnoses 90% of the time.” The Lancet abstract and the preprint are narrower: AMIE’s differential included the final diagnosis from chart review eight weeks later in 90% of cases, with 75% top-3 accuracy in the preprint. Blinded raters in the preprint did not find a significant difference between AMIE and the PCP on overall differential quality (p = 0.6) or on management-plan appropriateness and safety (p = 0.1 and 1.0). PCPs scored higher on practicality (p = 0.003) and cost-effectiveness (p = 0.004) of management plans. Those p-values are the authors’.
A list that contains the later chart diagnosis is not the same as a correct, complete, or safe explanation to a patient. For how to check a model’s medical answer, see verifying AI answers rather than treating 90% inclusion as a clinic-ready accuracy rate.
How did patients and doctors respond?
The Lancet abstract says clinical evaluators rated AMIE favourably on 87–100% of 17 criteria, and patients on 48–96% of 16 criteria. Patient attitudes toward AI improved after the chat and stayed higher after the physician visit. BIDMC quotes Adam Rodman, director of AI programs at the Shapiro Institute and a visiting researcher at Google during part of the study, on the point of the work: safety and acceptability in a real workflow, before anyone treats the bot as a workload tool.
PCP numbers are smaller than the 98 completed pairs. The abstract: 60 of 98 clinicians returned a post-survey; 44 had reviewed the AMIE transcript before the visit; 33 of those 44 found it helpful for preparation; 25 of 44 said it might have changed their behaviour. BIDMC adds that in one case a clinician called the interaction somewhat harmful because AMIE had included lymphoma among possible diagnoses and the physician worried the patient became anxious. Google’s blog compresses the 33/44 and 25/44 figures into “75% of cases” and “more than half,” without saying those denominators are the reviewed subset.
Patients in the BIDMC note still flagged confidentiality and whether they could trust the chatbot’s honesty. That is a different problem from a made-up citation. For invented medical claims, see how hallucinations show up in answers. The one logged hallucination in this study is a supervisor observation, not a measured hallucination rate.
What does this study not establish?
It does not show that AMIE shortens visits, reduces errors, or improves outcomes. Rodman and Marc L. Cohen, BIDMC’s clinical chief of primary care and a co-senior author, both say the design is feasibility and safety, not a workload or outcomes trial. There is no control arm. Every chat had a live physician on the line. The population is English-speaking adults with a single new complaint at one academic clinic. Psychiatric and pregnancy visits were excluded.
Funding and roles are disclosed. BIDMC says Alphabet funded the study and that Rodman was a visiting researcher at Google. ClinicalTrials.gov lists BIDMC as lead sponsor and Google LLC as collaborator. Google’s blog headline — that the study “suggests AI could improve patient-physician relationships” — is the company’s reading. The paper’s own limit is larger trials.
AMIE is also not a consumer Google product in this write-up. The blog calls it “our research diagnostic AI chatbot.” We have not seen a public endpoint, a price, or a clinical-use license. The inspectable 8 October facts are the Lancet abstract, the BIDMC note, the Google blog, the trial record, and the earlier preprint’s methods — plus the paywall on the journal’s full text.
Common questions
Can a patient open AMIE in Gemini today?
Not from these pages. Google describes AMIE as a research system used inside a supervised BIDMC study. There is no consumer download or clinic rollout in the notes we opened.
Does “no safety stops” mean the chatbot was unsupervised?
No. A physician watched every chat in real time and, per the registry, corrected errors afterward. Supervisors still logged one hallucination and five clarifications.
Why do some pages say 100 patients and others say 114 and 98?
The March preprint and the ClinicalTrials.gov enrollment field use 100. The Lancet abstract and BIDMC’s 8 October note use 114 enrolled and 98 completed. We are reporting both and dating the journal version to 8 October.
What to remember
Date the news to The Lancet and BIDMC on 8 October. Keep the 90% line as differential-inclusion on chart review, the 75% and 57% lines as a 44-clinician subset, and Google’s relationship headline as commentary on a funded feasibility study.
Sources & further reading
- BIDMC Researchers Conduct First Real-World Study of Safety and Quality of Patient-Facing AI in Primary Care ↗
- Study in The Lancet suggests AI could improve patient-physician relationships ↗
- Conversational diagnostic artificial intelligence in ambulatory primary care: a prospective feasibility study ↗
- NCT06911398: AMIE's Clinical Conversational Abilities in an Urgent Care Setting ↗
- A prospective clinical feasibility study of a conversational diagnostic AI in an ambulatory primary care clinic ↗
How this story was made
Written by Kristian Kostov with AI assistance and checked against the linked sources. Company performance claims are attributed to the company. Analysis reflects AiLookout’s interpretation; we have not independently tested the products discussed. Cover photography is illustrative and does not depict the specific announcement or product.
Our editorial standards





