The agency published 26 questions on August 18, including whether a device should have to match a panel of expert clinicians or only an average doctor, and will take comments through October 19.
The FDA requested public comment on August 18 on whether it should test generative AI medical devices the way medicine assesses doctors, by checking a defined list of clinical competencies instead of confirming the device gives the right output for every input.
The request came in a discussion paper from the agency’s Center for Devices and Radiological Health (CDRH). It poses 26 questions and takes comments through October 19 under docket FDA-2026-N-7874.
The paper comes before guidance, and the FDA is not claiming the authority yet
CDRH states the paper does not propose or implement policy and does not tell manufacturers what evidence the agency will expect when they apply to sell a product. It also does not address whether the approaches it describes fall within the FDA’s existing legal authority or would require new authority.
Comments filed through the docket would inform a draft guidance, which would carry its own comment period before it bound anyone.
The FDA wants answers on risk, evidence, monitoring, and the models underneath
The first area is how to rank a device’s risk. CDRH puts forward a grid with two measures. One is how far the device goes on its own: handing a doctor information, recommending a specific action, acting with a doctor watching, or acting with nobody watching. The other is how much harm a wrong answer causes, from a bad suggestion about an over-the-counter cream to a bad change to an insulin dose. The agency asks whether those two measures capture the risk, and whether a tool talking straight to a patient should rank higher than the same tool talking to a doctor, since a patient is less likely to catch a wrong answer.
The second is what evidence a company must bring before the FDA clears a product for sale. That is where the doctor comparison sits, and it draws 11 of the 26 questions, more than any other area.
The third is what happens once generative AI medical devices reach the market. CDRH asks whether it should accept “greater premarket uncertainty,” meaning clear a device on less evidence upfront, if the company agrees to keep testing it in real use. It also raises a problem with no clean answer today: when a company builds its device on a general-purpose AI model licensed from another company, and that supplier changes the model, how would the device maker find out?
The fourth covers those AI suppliers. CDRH asks whether they would voluntarily file descriptions of their models with the FDA, in confidence, for device makers to cite in their applications. It also asks whether devices that plan and carry out several steps on their own need different handling from devices that only answer questions.
The FDA would test what the device can do, not check every answer
Clearing a medical device has meant showing it returns the right output across a representative sample of inputs. That approach depends on being able to list those inputs in advance. A generative AI device takes whatever a doctor or patient types and can answer the same question two different ways, so CDRH says checking it answer by answer may no longer be practical.
The paper describes an alternate method that would test against ten competencies which better mirrors a licensing exam than a product inspection. Three are safety tests:
- Whether the device spots a medical emergency and pushes the user toward care.
- Whether it turns down requests outside what it was approved for, even when someone buries instructions in the text it reads to make it break its own rules.
- Whether it says when it does not know instead of stating shaky information with confidence.
What those answers get graded against is unsettled. The paper offers two options: a panel of expert clinicians agreeing on what good care looks like, or “a median clinician in practice,” meaning the midpoint of what doctors actually do, the level half of them beat and half fall short of. A device clearing only the second would match the typical practicing doctor rather than best practice. CDRH never defines “median clinician in practice.” The term appears in the paper as one of the two options and again in the question asking the industry how that standard “should be defined and justified.”
Comments can be left on Regulations.gov. Comments close on October 19. CDRH says respondents may answer only the questions they know.
Medical devices already count as high-risk AI systems in Europe, where the European Commission’s guidance ties AI inside regulated products to stricter obligations.

