Evaluating LLMs for Health & Wellness Information with the NYC Department of Health and Mental Hygiene
Written by Crescentia Jung
The Challenge: Assessing AI-Generated Health Information
Increasingly, the public is using large language models to find health and wellness information. Specifically, recent Pew surveys have found that 22 percent of Americans use chatbots for health-related information¹. In New York specifically, as many as 67 percent have used chatbots². However, it is not always clear which models are most appropriate for public-facing health use or where important limitations remain. Their responses may be inaccurate, incomplete, unclear, or inconsistent. Even a response that includes correct information may leave out important context or present risks in ways that affect how a user understands the answer.
Public health organizations, like the New York City Department of Health and Mental Hygiene, therefore need practical ways to assess how these tools perform and how they may be used responsibly.
This raises a broader question: How can public health organizations evaluate the strengths and limitations of AI-generated health and wellness information?
The Project: Comparing Consumer-Facing Large Language Models
As a Siegel PiTech PhD Impact Fellow with the New York City Department of Health and Mental Hygiene, I worked with public health professionals to compare four commercially available large language models: Gemini, Claude, ChatGPT, and Grok.
Our goal was to understand how the models performed when responding to health-related prompts. We examined not only whether their responses were correct, but also whether they were safe, clear, complete, and consistent across different ways of asking the same question.
Together, we:
Reviewed prior research on large language models and health information
Developed prompts and response-collection procedures
Created a blinded physician-rating process
Defined criteria for accuracy, public health safety, communication clarity, completeness, and consistency
Developed rating scales and calibration materials for physician raters
Each step required interdisciplinary collaboration. For example, we discussed when technically accurate information could still be communicated in a misleading way, and whether a change in tone also changed the substance of an answer.
My background in human-computer interaction helped me consider how users might interpret and act on a model’s response. The physicians brought expertise in clinical evidence, health communication, and the potential consequences of inaccurate or incomplete information. Combining these perspectives allowed us to develop an evaluation approach grounded in both research and public health practice.
Impact & Path Forward
Crescentia Jung
Ph.D. Student, Information Science, Cornell University
We are currently analyzing the responses and will use the findings to develop a research paper and a public-facing summary for the New York City Department of Health and Mental Hygiene.
The project is intended to help DOHMH better understand how generative AI tools may be used responsibly in health information settings and what guidance may be needed as these technologies become more common in everyday information seeking.
The evaluation framework may also be adapted to other health topics or used to reassess models as they change. Because commercial AI systems are frequently updated, their performance should not be treated as fixed.
This fellowship allowed me to apply my human-computer interaction research experience in a public health setting and learn from collaborators with different expertise. It reinforced that public-interest technology is not only about creating new tools. It also involves evaluating the tools people already use and developing evidence-based guidance for their responsible use.