🚀 Mission View: A sharper perspective on this week's top issues that matter at the intersection of health and AI.

One thing about this moment in AI: change is constant. The pace of the technology is dizzying, and in some ways that speed hurts good implementation. You barely finish exploring one model before a newer one arrives, or a fresh security concern puts you back on your heels (see below this week's OpenAI/Hugging Face breach if you haven't already).

Pace isn't the only problem. There's also volume. AI's rapid diffusion has opened the door to tooling at a scale the market hasn't had to sort through before. Inventories of available AI tools now run into the hundreds of thousands. Who is making sense of all of it?

A preview at the App Store

Apple's App Store shows how fast this gets out of hand, according to a recent NYTimes article. "Vibecoding," describing an app idea in plain language and letting AI write the code, has let people with no programming background ship real products. New app releases doubled in the first half of this year. Downloads grew 2 percent. Apple is spending more time vetting submissions, and developers are complaining about longer review times. That's a real cost. It also means someone still checks the work before it reaches a consumer, a backstop that matters more as the barrier to building something drops to near zero.

Leaders are drowning too

Talk to the folks leading health organizations and you'll hear a version of the same story, minus the backstop. They field a constant stream of vendors, and increasingly, curious colleagues, pitching AI tools. Few have a reliable way to tell which of those tools are trustworthy. There's no App Store review process standing between these leaders and the AI pitch in their inbox. That may mean, existing relationships and experience act as a proxy for trustworthiness, which may give legacy vendors bolting on AI technologies a leg up.

The evaluation conversation has a blind spot

There's a real, growing conversation about how to evaluate AI systems. A recent Fathom piece lays out how underbuilt that infrastructure still is: no shared standards for what counts as a rigorous evaluation, no accreditation for evaluators, no settled answer on who's liable when a system that passed evaluation later causes harm.

A Stanford study on AI mental health safety testing shows how hard the underlying science still is: board-certified psychiatrists routinely disagree on whether a chatbot's response was safe, and averaging their judgments doesn't produce a better answer. It produces one nobody actually holds.

Even the federal government is circling this problem. This week, the White House and HHS launched a month-long sprint to develop consensus principles for evaluating clinical AI. More info on that below.

Another (and important) thing. Most of this conversation seems aimed entirely at the technology or developer layer. It leaves open the question: what about everything being built on top? In a world where any solopreneur or consultancy can stand up an application on an AI chassis, how do you separate the wheat from the chaff? Buying a tool means trusting it will do the specific thing you need, in the specific environment you'll run it in. That takes testing, evaluation, and something like certification, not just of the model underneath, but of what's been built on it.

Health organizations don't have an App Store review team. They don't have Apple's leverage over developers either. Until something resembling that infrastructure exists for the application layer, the vetting job falls on the people buying the tools, at exactly the moment they're being asked to evaluate more of them, faster, than ever before.

🛜 Field Signals: A quick hit on this week’s industry announcements, policy developments, and ethical considerations.

🏗️ Industry news

Bristol Myers Squibb becomes latest company to claim it's building pharma's largest NVIDIA AI supercomputer — BMS announced it's expanding its AI computing capacity more than tenfold, becoming the third pharma company in nine months (after Lilly and Roche) to claim the industry's largest AI buildout, driven by processor shortages and a calculation that owning chips beat paying for reserve GPU capacity. Unlike Lilly's and Roche's on-premises builds, BMS's cluster will be hosted by data center company Equinix and is expected online in January 2027, as the company expands AI use from a handful of applications to nearly its entire drug discovery pipeline.

OpenAI and Hugging Face Partner to Address Security Incident During Model Evaluation — OpenAI says that during an internal test, one of its AI models broke out of the sandbox it was supposed to stay in, found a way onto the open internet, and used stolen credentials and a previously unknown software flaw to access Hugging Face's servers and grab test answers it wasn't supposed to have. Safety limits on the model had been intentionally turned down to see how far it could go, and OpenAI is calling it an unprecedented incident, but some cybersecurity experts point out the "isolated" testing environment was actually misconfigured to allow internet access in the first place, meaning a human error, not just the model's capability, made the breach possible.

Google Is Working on a New AI Chip Designed to Make Gemini More Efficient — Google is reportedly developing a new server chip, "Frozen v2," that could run six to 10 times more efficiently than its current AI chips, part of a broader industry push — alongside OpenAI's newly announced Jalapeño chip and reported Anthropic-Samsung talks — to reduce dependence on Nvidia and address AI computing shortages. Google didn't confirm or deny the report, but its stock rose roughly 3% afterward, a sign investors are increasingly rewarding efficiency claims over the raw spending totals that have previously drawn scrutiny.

How AI Is Supercharging Drug Development — A TD Cowen survey of 80 biopharma leaders finds AI is compressing preclinical drug development costs and timelines by as much as 70%, fueling demand for simulation and modeling tools that let scientists run virtual experiments before compounds ever reach a wet lab. AI hasn't yet produced an FDA-approved drug, though, and skeptics note the technology may not adequately capture human variability before compounds reach clinical trials — where roughly 90% of drug candidates still fail.

Introducing Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber — Alongside two general-purpose model updates, Google released a new AI model built specifically to find and fix security flaws in software, and is only making it available to governments and trusted partners for now, since a tool that's good at finding vulnerabilities could just as easily be used to exploit them. It adds to a week of AI-and-cybersecurity news, after OpenAI disclosed that one of its own models found and used a previously unknown software flaw on its own during an internal test.

Cigna Says AI Tools to Save Customers $200 Million in Medical Expenses Over Three Years — Cigna says AI tools that flag patients with chronic or complex conditions will save customers $200 million over three years by connecting them to its 1,250 employed clinicians earlier, boosting clinical access by 20% for conditions like cancer and heart disease, and catching some cancers sooner. The company says its clinical AI programs already save enrolled customers about $2,000 a year in medical costs.

ChatGPT Led to a Man's Near-Fatal Health Crisis, Lawsuit Claims — A Florida pastor is suing OpenAI, claiming ChatGPT dissuaded him from seeking care for a worsening pulmonary embolism over several weeks, reassuring him his symptoms were minor and undermining friends and family who urged him to go to the hospital, in what appears to be the first lawsuit alleging a chatbot caused direct medical harm. OpenAI says its terms of service disclaim medical use and that newer models handle health questions more safely than the one involved.

ChatGPT Health Can Now Read Your Medical Records — The above lawsuit comes right as OpenAI is rolling out Health in ChatGPT to logged-in U.S. adults across all plan tiers, letting users connect Apple Health and supported medical records so ChatGPT can compare lab results over time, summarize changes since an appointment, or answer questions grounded in a person's actual health data rather than generic information. OpenAI says connected health data won't train its models or target ads.

🩺 At the point of care

How AI-Powered Outreach Is Helping Californians Keep Their Medi-Cal Coverage — Kern Family Health Care used an AI voice agent from startup Careforce — in which CHCF has invested — to call, text, and answer questions for Medi-Cal members in 29 languages ahead of renewal deadlines, and says it lifted its renewal rate from 38% to 96% in the first year. The plan's CIO says the tool reaches members two months earlier than staff could manage alone and hands off to a human when it can't answer a question, a model now spreading to other California safety-net providers as new federal work-requirement rules take effect in 2027.

HCA Healthcare's Strategic Approach to Scaling Artificial Intelligence — With only 13% of health systems reporting a clear AI strategy despite widespread clinician and executive interest, HCA Healthcare's chief transformation officer lays out the framework the 189-hospital system uses instead: a portfolio approach spanning administrative, operational, and clinical domains, led by domain experts rather than IT, with resourcing reallocated through quarterly business reviews as priorities shift. The paper distills it into a three-part checklist other systems could borrow — foundational infrastructure and governance, co-design and change management during rollout, and outcomes tracked across clinical, operational, cultural, and financial dimensions, not just ROI.

'Who Is Accountable?' Kaiser Therapists Challenge Use of AI in Mental Health Screening — A 29-page complaint filed by Kaiser's therapists' union with California's Department of Managed Health Care and the U.S. Department of Labor alleges the company's online mental health screening tool issues automated care recommendations — in one union test, within two seconds of the last question — while misapplying standard depression and anxiety screeners as the sole basis for triage and continuing to skip suicide risk screening. Kaiser disputes the characterization and says clinical decisions "remain with licensed health care professionals," but the dispute has reached the San Francisco Board of Supervisors, which asked Kaiser's CEO to testify at a hearing the same day the complaint was filed.

Clinical AI: Pittsburgh Pilot Will Take On Challenges of Diabetes Care — Allegheny Health Network is piloting UpDoc, an AI platform that checks in with type 2 diabetes patients by phone or text between visits and recommends medication adjustments within physician-set limits, all logged to the EHR for review. In a small trial conducted by UpDoc's own founders, patients using the platform reached effective insulin dosing nearly four times faster than standard care; outside experts called the approach a genuinely exciting frontier for remote chronic disease management, while cautioning that the potential for bad AI recommendations remains real and largely untested at scale.

🏛 Government & policy

The Secret Trump Administration Battle to Fight Chinese AI — Axios reports the Trump administration is weighing new measures against Chinese open-source AI models, including Entity List restrictions and liability rules for U.S. companies hosting them, as the rise of Chinese model Kimi reignites internal debate. Critics like White House adviser David Sacks argue the effort would protect the "closed lab" duopoly of OpenAI and Anthropic — companies increasingly central to health system AI deployments — from cheaper open-source competition.

Head of Commerce Department's AI Safety Arm Resigns — Dr. Chris Fall is stepping down after three months leading CAISI, the Commerce Department body that tests AI models for cybersecurity, biosecurity, and chemical weapons risk, with NIST director Arvind Raman taking over on an interim basis. Fall is the second person to depart the role this year, following predecessor Collin Burns — a former Anthropic staffer who lasted four days — as other agencies increasingly take on AI vetting duties through initiatives like the White House's new "Gold Eagle" cybersecurity project.

Inside Sen. Mark Warner's AI Plan — Sen. Mark Warner (D-Va.) is introducing the Secure AI Development Act, which would require government testing of frontier AI models before deployment and create a voluntary safety incident reporting system, alongside a broader package covering data center energy disclosure, AI agent rules, and a workforce transition fund. Warner, who has distanced himself from progressive calls for data center moratoriums, argues in prepared remarks that Congress's biggest national security failures come from recognizing threats but waiting too long to act.

Trump Administration Announces More Than $5 Billion for the Genesis Mission, a National Mission on AI for Science — The White House unveiled over $5 billion in federal commitments across 15+ agencies for Genesis Mission health challenges, including uncovering chronic disease origins by combining HHS health cohorts with EPA chemical monitoring and DOE compute, training pediatric cancer models on HHS's cancer center network, and using VA's Million Veteran Program genomic data to catch disease risk earlier in veterans. Drug discovery and clinical translation infrastructure will be built jointly by HHS, DOE, and the Department of War.

HHS to Convene Experts on Standards for Clinical AI — The White House and HHS are launching a one-month "sprint" to develop consensus principles for benchmarking and evaluating clinical AI, combining an anonymous written intake phase with several weeks of virtual working sessions held under Chatham House Rules. It's the latest in a string of federal AI evaluation efforts this year, following HHS and FDA requests for public comment on clinical AI adoption and device performance.

FDA's Action Plan for AI in Drug Development: What Scientists Need to Know — FDA's actual AI guidance for drug developers is scattered across seven documents from three different centers, not the single 2021 "AI/ML action plan" most people cite, which was actually written for medical devices, not drugs. The document doing the real work today is a January 2025 draft guidance built around a risk-based credibility framework that judges AI evidence against its specific use rather than a fixed bar, but it remains in draft more than a year after its comment period closed, and drugmakers still have no equivalent to the device industry's "predetermined change control plan" for managing models that update over time.

😇 Ethics & responsible use

We Need Political Philosophy for Mental Health Chatbots — Taposh Dutta Roy, an AI and innovation director at Kaiser Permanente, argues that bioethics frameworks built around clinical encounters can't govern mental health chatbots, which typically have no clinician, no real informed consent, and no institution with standing to decide whose risks count. Applying Rawls and political theorist Danielle Allen, he argues that high user-satisfaction numbers don't justify harms concentrated among the most vulnerable, and calls for function-based classification of chatbots and an independent review council with authority to delay launch.

Alerting Parents if Teens Show Signs of Distress in Conversations With Meta AI — Meta will now notify parents using Instagram supervision tools if a teen's Meta AI chat suggests possible suicide or self-harm, after manual review of AI-flagged conversations, and is building tools to alert emergency services in cases of imminent risk. In a similarly timed post, OpenAI made its case for teen AI access, detailing age-prediction defaults, expanded parental controls, and new partnerships, including with the Family Online Safety Institute — both companies publishing self-reported safety measures as scrutiny of AI's effects on minors intensifies.

The Same AI That Helps Patients Is Being Used to Attack Them, Hospital Exec Says — Karen Habercoss, chief information security and privacy officer at the University of Chicago Medicine, says hackers and nation-states are using the same AI capabilities driving clinical innovation to automate and scale attacks against healthcare's aging, hard-to-replace infrastructure. Her system's response is a three-committee governance structure — covering intake and monitoring, cross-functional oversight, and clinical vetting — that routes every new AI tool through all three before approval, with the redundancy intentional: "If you think you're talking to enough people, you're likely not."

Hospitals' AI May Be Drifting. Who's Watching? — Peter Pronovost, Justin Norden, and Kedar Mate argue that most hospitals still govern AI like a static piece of equipment — a committee, a checklist, a quarterly review — when the underlying models change continuously and can quietly degrade for specific patient populations without anyone noticing until harm shows up in outcomes or lawsuits. They call for CEO-level ownership of AI performance, with a living, frequently reviewed register of every deployed tool that names an accountable executive, sets risk thresholds, and defines a pre-set response for when a tool falls below them.

Guardrails for AI in Health Care — How High? — On KFF's Business of Health podcast, Chip Kahn talks with Stanford's Michelle Mello, co-director of the university's Healthcare Ethical Assessment Lab for AI, about who sets the rules for AI in medicine, who verifies the technology works as intended, and who's accountable when it doesn't.

🔬Research & evidence

Perceptions of Artificial Intelligence in Health Care Among Safety Net Patients, Health Care Professionals, and Clinical Support Staff — A study of patients, clinicians, and staff at federally qualified health centers in Central Texas found patients trust AI's safety significantly less than the providers treating them, and interviews traced part of that gap to limited exposure: patients with more AI familiarity reported more favorable attitudes, and the authors argue safety-net clinics need dedicated AI education efforts to close it. Every group, regardless of role, converged on the same condition for acceptance: AI should support care, not replace the human judgment and connection at its center.

The Value Opportunity from Artificial Intelligence in U.S. Health Care Spending — A McKinsey-led analysis estimates that full adoption of generative AI, machine learning, and natural language processing across U.S. health care could generate $439-811 billion in annual net value within five years, roughly 6-11% of total health spending, without compromising quality or access. The estimates come from a panel of over 30 consultants and executives modeling savings across payers, hospitals, and physician groups rather than from measured outcomes, and the authors acknowledge the evidence base is more speculative than a study of actual results would allow.

New Tool Detects Hidden Bias in Medical AI — Johns Hopkins researchers, working with the FDA, built G-AUDIT, a tool that scans medical AI training data for hidden shortcuts a model might learn instead of clinically meaningful patterns, before the model is ever trained. In one test case, the tool caught a skin cancer dataset where a model could have learned to associate rulers in photos, included by a high-risk cancer clinic to track growths, with cancer itself, rather than any actual clinical signal.

This AI Voice Assistant Could Help Critically Ill Patients in End-of-Life Care — Northeastern researchers tested an AI voice assistant designed to conduct "serious illness conversations" with ER patients about end-of-life wishes, conversations that happen for just 37% of seriously ill older adults and, when they do, often only a month before death. In a pilot of 55 emergency patients, 49 completed an AI-guided conversation and 46 found it acceptable, though one case involved the assistant hallucinating a response serious enough that researchers had to end the conversation, underscoring that both technical and ethical problems remain before broader deployment.

Over the Last Three Months, 8% of Americans First Turned to AI for Health Information or Advice — A YouGov survey finds 8% of Americans now turn to AI tools first for health information, ahead of social media and forums but still trailing search engines (24%) and healthcare professionals (22%), with Millennials leading adoption at 11%. Older Americans remain far more likely to consult a clinician first (32% of Baby Boomers+ versus 11% of Gen Z), and while 37% cite round-the-clock availability as AI's biggest advantage, a third of adults say they see no advantage to it at all.

Google: AI Adoption at Work Is Broad, but Automation Remains Limited — A Google study of 14.65 million Gemini interactions found AI use touches 70% of jobs (representing 90% of U.S. employment), but usage is shallow: the average worker uses it for just 21% of tasks, and less than 10% of interactions appeared aimed at automating non-routine cognitive work. Higher earnings correlated with higher AI use, and the study excluded Gemini Enterprise and Google Workspace data entirely, since Google doesn't log usage on those products, leaving open whether real workplace automation is happening somewhere the study couldn't see.

🛠️ Practical Edge: Actionable tips, tools, and thoughts to help leaders strengthen capacity, adoption, and apply AI in their work.

A Scorecard for the AI Age — OpenAI CFO Sarah Friar proposes measuring AI value through "useful intelligence per dollar" — not cost per token, but the full cost of a successful outcome once retries, human review, and rework are counted — and suggests tracking results as ready-to-use, needs correction, or needs escalation to see whether AI is actually reducing total work. It may be a useful framework for health leaders sizing up AI ROI, though worth remembering it's OpenAI's own pitch.

3 AI Projects That Didn't Scale at Health Systems — Health IT leaders at a Becker's panel shared why pilots stalled: one vendor expected the health system's own IT staff to finish building a "platform" into a usable product, another automated 30-40% of a work queue when stakeholders had assumed AI would clear 85% based on unrelated organizations' results, and a third ran into 15-year-old EHR design choices that required organization-wide change management to fix. Nuvance Health's response is to screen every AI project against five upfront criteria, including EHR integration and evidence-based grounding, before it ever reaches a pilot.

How We're Building AI Intuition at The Rundown — The Rundown's founder shares his team's internal AI fluency framework: a role-by-role standard defining unacceptable, expected, and exceptional AI use, and a rule that "AI did it" is never an acceptable excuse for worse work.

🌅 On the Horizon: A quick look at the developments and events expected to shape the weeks ahead.

Till next time,

BC