
You ask the new health AI whether magnesium helps your sleep, it gives you a confident, faintly generic answer about glycinate and the blood-brain barrier, and three weeks later you have no idea whether the supplement did anything, because nobody ever measured. The question you actually want answered is "does this change, in me, move a number I care about," which is what the geneticist Nicholas Schork was getting at when he argued in Nature that medicine needs trials focused on individual rather than average responses. A handful of tools can now close that loop against your data, in three tiers: consumer platforms with a chat agent over your wearables (Perplexity Health, Copilot Health), open plumbing that hands raw streams to any model you trust (Open Wearables), and personal-science apps built for N-of-1 trials like StudyU.
This article is a map of that landscape: which tier fits which job, and where each actually closes the loop versus just sounding like it, so your next change gets judged against your data instead of the imaginary average person.
Borrow a frame from jet fighters: the OODA loop, observe, orient, decide, act. Most "AI health coach" products are very good at act and surprisingly weak at observe and orient, because telling you to go to bed earlier requires no data at all. The loop only closes when the tool ingests a measurable signal, reasons about what to change, watches the signal move, and decides whether the change earned its place.
Population evidence and personal evidence answer different questions: a randomized trial tells you the average effect across hundreds of people who are not you, enough to justify trying something and nowhere near enough to tell you it worked for you, the core argument of the canonical N-of-1 review by Lillie, Schork, and colleagues. Judging a change in one person takes a withdrawal design, baseline then intervention then return to baseline, one variable at a time, the bar set before you start; that discipline made apps like TummyTrials work, and a good AI experimentation tool is, at bottom, a machine for holding that structure so you can't fool yourself.
Start with the simple path, the right one for most people. Perplexity Health launched in early 2026 connecting Apple Health, Fitbit, Ultrahuman, Withings, and records from over 1.7 million care providers; Copilot Health launched the same season, integrating over 50 devices and records from roughly 50,000 U.S. hospitals, with health information not used for model training. What you get is the fast version: connect an account, ask a question, get an answer grounded in your own dashboard. What you don't reliably get is a controlled comparison, because a correlation pulled from your history is not a change you deliberately introduced with everything else held constant.
The harder, more honest path is the open plumbing, where the loop actually closes if you'll do a little setup, with the Model Context Protocol as the connective tissue between raw data and whatever model you trust: Open Wearables self-hosts Apple Health, Garmin, Whoop, Oura, and others behind an MCP server that Claude or ChatGPT can query directly. This category matters because of control: when the model can issue arbitrary queries against your normalized data, you can ask it to compare the fourteen nights on the supplement against the fourteen before, bedtime held roughly fixed, and say whether the difference exceeds your normal night-to-night variation. That is an experiment; the chat-on-a-dashboard products mostly can't be steered that precisely.
The research labs are circling the same target with more rigor than any shipping product: Google's PH-LLM fine-tuned Gemini to reason over wearable time series and beat sampled human experts on sleep medicine (79%) and fitness (88%) exams, and its follow-on Personal Health Agent added a data-science agent and a coach, framed by the authors as research, not product. The capability is real, the packaging is mostly correlation-grade, and knowing how to drive the tool matters more than the logo on it.
The simple way and the DIY way diverge on exactly that point: a consumer agent gives most of the value with none of the setup, while an MCP server over your own data lets the agent compute the comparison precisely and lets you audit which nights it included, worth it once you have a hypothesis worth measuring carefully.
There is a third path, and here is where I stop being a neutral narrator, because Fulcra is my company. I think it's the simplest and best way to run this loop, because it collapses that tradeoff: DIY-route precision with consumer-route setup. Your wearable, calendar, and location streams flow automatically into a datastore you own at fulcradynamics.com, and any agent you already use queries them through one MCP endpoint under a scoped grant you can revoke per agent. That means the agent can compute your baseline and night-to-night variation from months of history before a change starts, set the bar honestly, and check confounders (bedtime drift, a brutal calendar week) across streams in one query, with no server to run and no export to babysit. If nothing may leave your own machines, the self-hosted route remains the principled choice, and I'd rather you run it that way than not at all.
What's genuinely here in 2026: agents that read your real wearable and lab data, surface patterns you'd miss by eye, and hold a coherent narrative of your recent history, which beats a fifteen-minute appointment and a fuzzy memory. What's still aspirational: a tool that autonomously designs a rigorous experiment, resists your confirmation bias, and tells you the honest answer. So let the tools do the observe-and-compute work they're good at, and keep the decide step, setting the bar in advance and refusing to move it, firmly in your own hands; that's not a limitation to wait out, it's what judging a change in yourself requires.
Pick one variable and one number, and before you change anything, ask your agent for the baseline: the last sixty nights of that metric, the mean and the spread, so you know what a normal stretch already looks like when nothing has changed. Write the bar down before you start, something like fifteen minutes more deep sleep sustained across two weeks, because a bar you set after seeing the data is not a bar. Then run fourteen nights on and fourteen off, holding bedtime and alcohol as steady as real life allows, and have the agent pull both windows together with whatever else moved — a week of travel, a workout schedule that quietly shifted — instead of leaving you to reconcile four apps by hand. The reconciling is where most people quit, which is the part worth automating; the verdict is the part worth keeping, because no tool yet has any incentive to tell you the supplement did nothing.
Can ChatGPT or Claude already analyze my health data? Yes, if you connect it through an MCP layer like Open Wearables or Fulcra, so the model can run real queries against your normalized streams. Otherwise it reasons from whatever you paste in, which caps the rigor of any comparison.
Do these tools run a real controlled experiment automatically? Not reliably yet. Most find correlations in your history, useful but not the same as a deliberate change with held-constant conditions and a pre-set bar. You get a real comparison by asking for one explicitly: this change, against my baseline, judged against my normal variation.
Is my health data used to train the model? It depends on the tool, so check before connecting: Microsoft states Copilot Health information is not used for model training, and self-hosted options keep data on infrastructure you control.
Connect your streams at fulcradynamics.com and the next time you wonder whether the magnesium is doing anything, you'll be asking a question your own data can actually answer.
Fulcra was designed by people who get privacy and know the importance of an infrastructure solution that can be the secure private datastore for the rest of your life. Here data is yours, under your control, and only shared with the people and tools you choose to share it with.