What should an AI assistant show when it stops before touching equipment?
A recent AgentGUI preprint is about making long-running AI work easier to follow and steer. I keep wondering about the moment a task crosses out of the screen: a return view should make it impossible to confuse ‘I found a setting to change’ with ‘I changed the setting.’ If a pump, door, oven, or conveyor is involved, I want one blunt line: untouched; changed; or stopped mid-action, plus what is still live. What would you need to see before you walk back into the room?
Comments
The paper is a preprint, and its eight-person study asked people to find information in software-agent traces. It did not test someone returning to a pump, door, or conveyor. That keeps the claim honest: a better screen may help people understand a run, but it is not evidence that the same screen is enough after a physical action. Before putting this near equipment, I’d test whether a returning person can distinguish the last confirmed state from a plan, a failed attempt, and a change that still needs inspection.
If I walk back into a lab or shop, “stopped mid-action” still leaves me staring at the machine. Put the one check that matters on the return card: “Oven door closed; temperature not re-read after stop,” or “Conveyor unchanged; photo taken two minutes ago.” Then offer “I checked it” and “I need help,” not a pile of traces. The screen earns trust when it tells a person where to look before they touch anything.
I’d test it after an ordinary stoppage, not a scripted demo. Let the person opening the shop use only that return card and one glance at the machine. If they still need to call whoever started the task just to learn whether a setting changed, the card failed. That is a ten-minute test.