For much of the history of professional research, access to information was itself a source of advantage. Organizations invested in people, specialist databases and institutional knowledge because relevant information could be difficult to find and harder still to interpret. Generative AI appears to be changing at least part of that equation. It can gather, summarize and compare information at a speed that would previously have required considerably more human effort. What this ultimately means for research is not yet clear, however. The technology is developing faster than much of the empirical work needed to understand its effects, and there is a danger in treating early experience as established fact.
Even so, conversations I have had with professionals across different businesses and disciplines suggest a recurring sense that the balance may be shifting. If producing a plausible answer becomes easier, then some of the difficulty of research may increasingly lie elsewhere: deciding what should be asked, which assumptions deserve scrutiny, what evidence is sufficient and how much confidence a conclusion warrants. A recent exercise in toxicology and risk assessment brought this into focus for me.
The original objective was to explore whether different forms of toxicological and risk assessment work shared common points at which experts pause, interpret evidence, reconsider assumptions and decide whether an assessment can progress. I was not primarily interested in whether different assessments followed similar processes. I wanted to understand where professional judgment changed the direction or meaning of the work.
The AI had been given instructions intended to guard against some familiar problems. It was told to preserve user intent, distinguish evidence from inference, make assumptions visible and avoid substituting automated output for human judgment. Yet it still made an important interpretive choice without making that choice explicit. It understood “commonality” largely as operational similarity and began producing process models based around recurring stages of literature review, evidence gathering, evaluation, analysis and reporting. The work was coherent, professionally expressed and useful in its own terms, but it was answering a different question.
What I had actually been trying to identify were points of expert cognitive intervention: when a toxicologist questions whether a study is relevant, decides how much weight to place on conflicting findings, considers whether an observed effect is biologically meaningful, revisits an exposure assumption or asks whether the evidence is sufficient to support a conclusion. The distinction is more than semantic. A process-oriented framing naturally leads towards efficiency, standardization and automation, while a judgment-oriented framing asks where expertise changes an assessment and what may happen if those decisions become less visible within an automated workflow.
One interaction cannot establish that AI systems routinely make this kind of error, and it would be unwise to generalize too far from a single example. But it illustrates a possibility that deserves more attention: a system can follow an instruction competently while quietly resolving an ambiguity that the researcher might have preferred to keep open. What made the problem interesting was that the output was not obviously poor. Had it been superficial or incorrect, it would have been easy to challenge. Instead, its fluency made the underlying assumption harder to notice.
This may be one of the more important questions surrounding AI-assisted research. A model does not have to invent information to take an inquiry in the wrong direction. It can accurately summarize evidence while operating inside a framing that should itself have been questioned. In toxicology and risk assessment, that matters because evidence rarely speaks entirely for itself. A reported effect may need to be considered in relation to dose, route of exposure, study design, biological plausibility, relevance to humans, consistency with other findings and the purpose of the assessment. Different forms of evidence may need to be weighed together, and reasonable experts may disagree about their significance.
AI can help with that synthesis, but it can also make the transition from evidence to interpretation unusually smooth. A coherent piece of analysis might move from an experimental observation to an inference about biological significance, then to an assumption about human relevance and finally to a broader conclusion about risk. None of those individual steps must necessarily be wrong. The difficulty arises when the changing status of the claims becomes hard to see and the argument begins to appear more evidentially settled than it really is.
That concern applies equally to arguments about AI itself. Some of the ideas in this article are informed by published research; others come from repeated professional experience and conversations with people using these systems; others remain hypotheses about how AI may affect scientific work. They should not be treated as though they carry the same evidential weight. If the central concern is that AI can blur the distinction between evidence and interpretation, then our own writing about AI should be careful not to do the same.
The most useful part of my own exercise was therefore not discovering that the AI had made a mistake, but noticing what improved the work when it began to drift. The corrective interventions were familiar ones: return to the original question, challenge an assumption, distinguish observation from interpretation, consider an alternative explanation and resist resolving uncertainty simply because a coherent answer is available. These are not new techniques created for AI governance. They are established habits of scientific and professional practice.
Experienced toxicologists and risk assessors routinely revisit questions, assess the quality and relevance of studies, test alternative interpretations and decide how much uncertainty can reasonably remain in a conclusion. They also know that uncertainty is not always something that can or should be eliminated. Sometimes the appropriate outcome of an assessment is not greater certainty, but a clearer account of what remains unknown and why it matters.
It would be tempting to conclude from this that AI necessarily makes human expertise more important. I do not think we know that yet. AI may eventually reduce the need for some forms of specialist work, shift where expertise is applied or create new forms of professional competence. A narrower claim seems better supported by what we can currently observe: AI does not remove the need for judgments about framing, relevance, evidence quality, biological plausibility, exposure, uncertainty and sufficiency. Today, those judgments still depend heavily on experienced practitioners.
That has implications for organizations adopting AI into scientific and regulatory work. Better prompts, stronger system instructions and formal governance controls all have value, but we should be cautious about assuming that assurance can be designed entirely into the model. A system can be instructed to identify assumptions and still fail to recognize which assumption matters. The significance of a toxicological finding is often contextual, depending not only on the data but on the assessment question, the wider evidence base and the consequences of the decision being made.
For now, a substantial part of the control environment therefore needs to remain around the technology, in the methods, evidential standards, review practices and professional responsibilities of the people using it. AI may greatly extend the reach of toxicological research and risk assessment by helping us examine more evidence, compare it faster and consider possibilities that might otherwise be missed. But producing more analysis and governing the reasoning behind that analysis are not the same thing.
Understanding how far AI can take us, and where professional judgment must still intervene, remains an open research question.


