Here is a strange test, and what came of it.
SuperTruth is a company that scores whether health data can be trusted. Its software reads a medical record and gives it a number from 0 to 100. That number is called the Data Trust Index™, or DTI™. High means the record holds up. Low means something is off.
This weekend the company handed that job to a fruit fly’s brain.
Not a live fly. A map of one. In June, scientists at HHMI Janelia Research Campus, the University of Cambridge, the MRC Laboratory of Molecular Biology, and Google Research finished mapping a fruit fly’s entire central nervous system- every nerve cell and every connection between them- by slicing the fly thin and photographing each slice under an electron microscope. They put the map online for anyone to use. It has 166,700 nerve cells and 6,242,118 connections. Those scientists were not part of this test and have not endorsed it.
SuperTruth took the map as is. It did not add a connection or move one. It only adjusted how strong each connection was, like turning a volume knob up or down. Then it trained that wiring to copy the DTI score on 20,000 medical records built for the test. No real patient data was used anywhere in this study.
Then it gave the same job to four AI models people know by name: Claude Opus 5, GPT-5, Grok 4, and Gemini 3 Flash. Each one got SuperTruth’s published paper explaining how the score works, plus 300 of the same test records. Each ran three times.
The fly’s wiring matched the score’s trust level 84 times out of 100. The four models matched it 20, 29, 28, and 45 times out of 100. Ask the fly twice, and you get the same answer twice. None of the four models did that on every record. SuperTruth asked every model for its most repeatable setting (temperature 0); two vendors’ services accept that setting and two do not in their reasoning mode, and the paper records which, so part of the models’ variation may come from the setting rather than the model.
Two things to be clear about. This test measured how closely each one matched SuperTruth’s own score. It did not decide who was right about the records. And the fly was trained on SuperTruth’s scores, while the models only read how the score works. So the company ran a second round and gave the models 100 scored examples to learn from. They improved a lot, from 54 to 76 out of 100. The fly still led at 84, and it gave the same answer every time it was asked.
“When a hospital acts on a bad record, a real person pays for it. That is who this is for. People have been told that bigger AI means smarter AI, and that a confident answer is a correct answer. Neither is true. What decides whether AI helps a patient or hurts one is the data it was handed and whether anyone checked it. A fruit fly’s brain just showed how much that matters,” said Bobby Hill, Co-Founder and Chief Executive Officer of SuperTruth. This study measured agreement with SuperTruth’s DTI score, not clinical correctness or patient outcomes, and it says nothing about any model’s safety for patients.
Now the part the paper is actually about. SuperTruth scrambled the fly’s wiring and ran the test again. The scrambled version did just as well. So did a random web of connections of the same size. So the fly’s exact wiring did not matter. What mattered was the kind of wiring: a fixed, thin web where every connection either pushes or pulls, and nothing gets rewired. A normal trained network with the same number of adjustable parts did worse. The paper calls this “the substrate, not the anatomy.”
The fly did not pass every bar SuperTruth set for it before the test began. It passed three of five. It missed the right trust level (87 percent on the full test set, against a goal of 90) and the part of the score that tracks how recent a record is. The company is saying so, not burying it.
“We publish the result either way. We wrote the controls and the decision rules before the first run, and a wiring that fails is a fact worth having,” said Jason Alan Snyder, Co-Founder of SuperTruth and the paper’s author. “Intelligence is structure, not scale. We borrowed a brain to prove it.”
“The dominant assumption in health AI is that a model’s judgment can be trusted because the vendor says so. That is not trust. It is deference,” Snyder said. “Trust stops being a vendor’s claim and becomes a property you can audit.”
How the test was run, in plain terms. The rules were written down and dated before the first run, with a promise to publish no matter what. The test records were kept apart from the training records and checked before any score was read. The four AI models were called through each company’s own service on September 20, 2026, all with the same instructions. Each company offers newer versions of its models, and a different version or different instructions could change the answers. Every call and every price is in the paper. All five test runs are done, and the paper’s third version holds them. The findings did not change as the seeds came in.
You can check it all yourself. The 20,000 records, the score for each one, and the trained fly model are free to download at https://doi.org/10.5281/zenodo.22865020. The paper, “Intelligence Is Structure, Not Scale: A Whole Central Nervous System Connectome as a Fixed Substrate for Scoring Health Data Trust,” is free at https://doi.org/10.5281/zenodo.22865214. A plain summary and a 3D view of the wiring are at https://supertruth.ai/research/connectome. SuperTruth has released everything from the study that doesn’t reveal its own inventions. If you want to know what was held back, ask, and the company will work with you. The paper describes a second test on the behavior of software agents. Those results are not out yet.
As far as SuperTruth knows, as of September 21, 2026, nobody has used the full wiring map of a nervous system to score whether health data can be trusted before. The paper describes how the company looked for earlier work.
SuperTruth’s scores rate data records and software-agent behavior. They do not diagnose or treat anyone, do not make recommendations about any patient, and are not meant for medical decisions.
The paper ends with a section the author marks as his own view, not a finding. “Simply put, the future of intelligence is analog,” Snyder writes.
Claude and Claude Opus are trademarks of Anthropic, PBC. GPT-5 is a product of OpenAI, Grok 4 of xAI, and Gemini 3 Flash of Google LLC; each name is the property of its owner. Anthropic, OpenAI, xAI and Google are named so readers can see what was tested. None is affiliated with SuperTruth and none has reviewed or endorsed this work.
- Paper: https://doi.org/10.5281/zenodo.22865214.
- Summary and 3D viewer: https://supertruth.ai/research/connectome.
- Data: https://doi.org/10.5281/zenodo.22865020.
- The DTI paper this work builds on: https://doi.org/10.5281/zenodo.19601616.
SuperTruth is a data truth company. Before an AI model acts on a record, a place, or another agent, we prove it.
It works in three parts. DTI™ is our proprietary trust score: 0 to 100, across eight dimensions, on any record. DataSpine is our sourced map of every place in the U.S.: tens of millions of data points, each with its source and vintage. VIGIL scores the AI agents that use that data and locks the result in a hash-chained ledger nobody can quietly edit.
We started in health because a wrong record there costs the most. The same proof works in finance, media, and anywhere you trust a machine to act.
SuperTruth: Data truth is AI truth.
The DTI methodology is published and citable: “The Data Trust Index: A Multidimensional Framework for Evaluating Health Data Integrity in AI Systems,” SuperTruth, Inc., 16 April 2026, DOI 10.5281/zenodo.19601616.
rheanna@supertruth.ai
SuperTruth