How We Use AI at Evidence First

0:00
-7:52

Evidence First. Conclusions Second. That is the idea behind everything we do.

The process starts with people. Humans decide which questions are worth asking, how those questions should be framed, and what we need to understand before reaching a conclusion. From there, we use AI to investigate those questions at a scale that would be impossible for a human research team to match in the time required to publish.

For most of history, research has had one unavoidable limit: human attention. There is only so much information one person, or even a large team, can find, read, compare, and make sense of in a reasonable amount of time.

We are not talking about reading a few articles and calling it research. Depending on the question, we may be working across thousands of sources and large datasets containing millions of individual observations. That is an enormous amount of information to search, compare, evaluate, and ultimately make sense of.

Consider the math. Imagine a researcher spent just one minute looking at each title and abstract to decide whether a source might be relevant. Screening 10,000 records would take about 167 hours. Screening 100,000 would take more than 1,600 hours. One million would take more than 16,000 hours, or roughly eight years of full-time work for one person.

And even one minute is nowhere near enough time to actually understand the evidence. It does not include reading the full study, understanding how the research was conducted, checking the numbers, evaluating its strengths and weaknesses, tracing claims back to original sources, comparing it with conflicting research, or figuring out what the entire body of evidence means.

That is where AI changes what is possible.

AI allows us to search across enormous amounts of information, quickly narrow that universe to the material most relevant to a question, analyze the important evidence in greater depth, compare findings across sources, work through large datasets, identify disagreements and gaps, and organize everything into a body of evidence our team can evaluate.

At this scale, the challenge is not simply finding information. It is making sense of an overwhelming amount of it. AI gives us the ability to work across a universe of evidence that a human team could not realistically read, compare, weigh, and synthesize quickly enough to publish.

The point is not simply to find more information. The point is to have a much better chance of finding the best information.

Not all evidence deserves the same weight. A carefully conducted study is different from an anecdote. An original document is different from someone else’s interpretation of it. A finding supported by many independent sources should generally carry more weight than a single surprising result that has never been confirmed. A dataset containing millions of observations can be powerful, but only if the data is relevant, collected properly, and interpreted correctly.

That principle applies whether we are investigating science, health, business, public policy, politics, history, or anything else.

We are not here to confirm what people already believe, tell an audience what it wants to hear, or shape a conclusion just to make one side happy. We are here to understand what is true as best the evidence allows.

That means applying the same standards to every claim, following the strongest evidence wherever it leads, and being willing to reach conclusions that may be inconvenient, unpopular, or different from what we expected. Sometimes the evidence strongly supports a conclusion. Sometimes it is genuinely mixed. Sometimes there simply is not enough good evidence to know.

Our job is to tell you which situation we are dealing with. Not to manufacture certainty. Not to soften a conclusion because someone may not like it.

Just as importantly, we do not ask AI to build a case for a conclusion we have already chosen. We ask it to challenge the conclusion as it develops. What evidence contradicts it? What are the strongest alternative explanations? What important sources might we be missing? What would have to be true for the conclusion to be wrong?

The goal is not to prove ourselves right. It is to make it difficult for a weak conclusion to survive.

Once the evidence has been gathered, compared, and synthesized, AI also creates the first draft. That matters because by this point we may be working with thousands of sources, potentially millions of individual data points, competing findings, qualifications, methodological differences, and layers of context.

That is far more information than a person could quickly hold in their head and turn into a coherent story. The challenge is not simply summarizing a few studies. It is taking a massive body of evidence and figuring out how to explain what it means to someone who may know nothing about the subject.

AI is especially useful here because it can work across that entire body of evidence, identify the most important ideas, connect findings across sources and datasets, and begin turning all of that complexity into a clear narrative. It is instructed to explain unfamiliar concepts in plain language, provide the context people need, use examples and analogies when they make difficult ideas easier to understand, and organize the evidence so the story is engaging and easy to follow.

At the same time, it has to preserve what matters. It should not make something sound certain when the evidence is uncertain. It should not erase disagreements that matter. It should not strip away important qualifications just to make a story simpler.

The goal of the first draft is not simply to make the research shorter. It is to take an enormous amount of complicated information and turn it into a clear, accurate, and engaging explanation of what we know, what we do not know, what the strongest evidence suggests, and why it matters.

But AI does not get the final word.

It can misunderstand a study, miss context, give too much weight to weak evidence, overlook something important, or simply be wrong. That is why the final stage is human.

Our editors go back to the important original sources, verify the evidence, challenge the AI’s synthesis, look for what may have been missed, make sure the draft tells the full story, and decide whether the conclusions actually follow from the evidence. Then they shape the piece so it is clear, engaging, accurate, and worth reading.

If the evidence is strong, we tell you it is strong. If it is weak, we tell you it is weak. If credible evidence points in different directions, we explain that. And if better evidence comes along later, our conclusions should change.

So the process is simple:

  • Humans choose the questions and decide what needs to be understood.

  • AI searches, analyzes, compares, and synthesizes the evidence, works across massive amounts of data, and creates the first draft.

  • Humans verify the evidence, challenge the conclusions, and shape the final story.

AI gives us scale. Humans provide judgment and accountability.

We are not here to tell people what they want to hear. We are here to get as close to the truth as the evidence allows and explain how we got there.

Because at Evidence First, the conclusion is never supposed to come first.

Evidence First. Conclusions Second.