What Faraday actually tested
Inherent calls its task set Replica. The initial suite contains 310 tasks drawn from 100 machine-learning and AI-for-science papers. Each task asks the agent to recreate a figure with a limited time and compute budget, without access to the original plot. That is a narrower claim than discovering new science, and the distinction matters.
A recreated plot can be a helpful checkpoint. It does not by itself establish that the interpretation is sound, that the underlying data were handled correctly, or that the result will survive a changed dataset or a new lab. Faraday’s launch page presents the work as a step toward an AI scientist; readers should keep it in that category: an early research-system result with a disclosed task design, not proof of an autonomous scientist.
The small-model angle is interesting, too. Inherent says Faraday uses 27 billion parameters while relying on coding agents as tools. If that approach holds up beyond its own suite, the cost of testing paper-replication assistants could drop. But the number that decides whether a research group keeps one will be the full review cost, not parameter count.
The missing deliverable is a working notebook
An AI research assistant should return more than a plot and a prose summary. A useful handoff names the paper and figure, data and code versions, assumptions it made, compute budget, tool outputs, failed attempts, and the checks that remain for a human. If a teammate cannot rerun the work or tell why the agent abandoned one path, the apparent time saving turns into an investigation.
That record also makes disagreement cheaper. A scientist may accept the reproduction but reject a preprocessing choice, a threshold, or a proxy metric. With the trail visible, they can change one thing and rerun it. Without it, they have to start from the finished chart and reconstruct the experiment backwards.
How to test an AI assistant for research this week
Pick five figures from work your team already understands. Give the assistant a fixed time and compute budget. For each attempt, record whether the figure was reproduced, how many human minutes it took to verify, whether the starting data and assumptions were named, how many failed routes were visible, and whether a second person could rerun the result without a meeting.
Compare that with the ordinary process, including the time spent finding files, explaining context, checking code and repairing a plausible-but-wrong result. A fast first chart is useful only when the next person can tell what happened. That is the bar Faraday and every other AI research teammate should have to clear.