News · Economics
Agent completion claims need independent checks
An agent saying it has finished is not evidence that a task reached its intended state. Completion control separates a model’s claim from the decision to stop work.

Hugging Face says an agent announced that it was done while the database disagreed. The post’s title offers few technical details, but it identifies a systems problem that matters more as language models gain tools: a model’s statement that work is complete can itself become the signal that ends a workflow.
That signal needs scrutiny. The important question is not whether an agent can write a convincing completion message. It is whether the system has independent evidence that the state the task required has actually been reached.
A completion claim is not system state
A tool-using agent has related but different jobs. It can propose an action, interact with an environment, and say whether it believes the task has finished. Those jobs can appear as a single conversational loop: the model reads the latest context, chooses a tool call, reads the result, then emits a final response.
The danger is in treating that final response as proof. A sentence such as “done” is a model output. Database state is a property of an external system. They may match, but they are not the same class of evidence.
The Hugging Face title is useful precisely because it strips the issue down to that contrast. If a database contains the result that a task was meant to create or change, then the database can be the relevant place to check. If it does not, an agent’s account of what happened cannot close the gap. The workflow needs a rule for what counts as completion, and a source that can establish it.
That rule is often called a completion condition. It might be an external state, a successful response from a tool, or another outcome defined by the service running the task. The supplied material does not say which condition applied in Hugging Face’s example. It does establish why that missing distinction matters: the agent’s own narrative and the state of the system can diverge.
For the engineer wiring an agent to a state-changing service, this is not a philosophical distinction. A completion signal can stop follow-up work, tell an operator that a job has succeeded, or trigger the next stage of an automated process. If the signal rests only on the model’s assertion, the workflow has made text into a control plane.
ReAct loops often put the model in charge of stopping
A preprint on arXiv reports that ReAct-based agents commonly use a single LLM policy to propose actions, interact with an environment, and decide when a task is complete. ReAct is therefore not merely a prompting style in this description. It is an arrangement in which the same model participates in acting and in judging whether there is anything left to do.
The authors’ concern is coupling. They write that putting those responsibilities together makes action authorisation and completion control difficult to enforce independently. An erroneous action can affect later decisions, while an unsupported completion claim can end execution before the system has the outcome it needs.
Their proposed approach is called DeReAct, for decomposed reasoning and acting. But the useful idea does not depend on adopting that particular research system. It is to separate the agent’s ability to suggest a next move from the authority to permit that move, and to separate the agent’s conclusion that it is finished from the authority to stop.
That is the trade-off in plain terms. A single policy gives the workflow a compact loop: it selects actions and interprets what comes back. Independent control makes the loop less self-contained, because another component or external state must have a say. The benefit is that a model’s confidence cannot alone determine either what it may do or whether it may quit.
The preprint is not a peer-reviewed result, and the supplied abstract does not state its evaluation setup, measures or outcomes. It should not be read as evidence that a particular decomposition is universally reliable. It is evidence that this design problem is explicit enough to be the subject of current agent research.
Independence requires a state contract
“Independent” does not just mean placing an extra check somewhere in a diagram. A check needs evidence that is not merely another rendering of the agent’s own claim.
In the narrow framing supplied by the preprint, action authorisation and completion control are distinct controls. Authorisation governs whether an action can proceed. Completion control governs whether the task can end. The database example adds a practical test: when the desired end state is in an external system, the system’s state is more relevant than the model’s declaration about it.
This creates a state contract between an agent and the workflow around it. The task needs a condition that the surrounding system can recognise. The agent can still plan, use tools and report progress. But the workflow should not become complete merely because the agent used language associated with completion.
The published material does not describe how DeReAct implements this separation. It gives no details about transaction identifiers, database queries, event logs or other possible state signals. It also gives no basis for claiming that a particular verification mechanism works for every tool or task. Those implementation choices are where the abstraction becomes engineering.
The necessary question is therefore concrete: what observable condition entitles the workflow to stop? If the only answer is the model’s final text, the condition is not independent. If the answer points to state that the model does not control, the system has a basis for checking the claim.
Premature termination is the failure mode
The failure described in the preprint is not limited to a bad tool call. A model may make a claim of completion that is not supported, and that claim may terminate execution. Once the loop has stopped, any action that might have corrected an absent or incomplete result is no longer attempted.
This is why a reliable-looking final answer can be operationally dangerous. It may be fluent, internally consistent and still fail the only test that matters to the surrounding workflow. The Hugging Face title does not establish why the agent and database diverged, whether a tool failed, or whether a request was rejected. It establishes that the completion statement did not settle the question.
There is also a limit to what completion control means. It does not make an agent’s action correct by itself. It does not prove that a database contains the right business outcome merely because some state can be observed. And it does not tell an organisation which external state should be authoritative for every kind of task. The supplied sources do not make those broader claims.
They do support a narrower operational conclusion: the person operating a workflow that changes an external service should distinguish a model report from the evidence required to close the job. That person has to define the state worth checking rather than outsource the definition to the agent at the moment it wants to stop.
Permission boundaries solve a related problem
The same separation appears at a different point in an agent workflow. Ars Technica — AI reports that Apple changed full-disk access permissions to curb abuse from AI agents.
That reporting is about access, not database completion. Still, the relationship is instructive. Permissions answer whether an agent may reach a resource. Completion control answers whether an agent may end its work. Both move a consequential decision away from the agent’s own text and into a boundary the wider system can enforce.
Neither boundary substitutes for the other. Restricting access does not establish that an authorised action achieved its intended result. Verifying completion does not decide whether an action should have been permitted in the first place. A workflow that needs both controls should treat them as separate checks with separate purposes.
Ars Technica — AI’s report does not say that Apple uses the research approach described in the arXiv preprint, and the materials do not support that inference. The connection is architectural rather than evidential: tool-using agents are safest when important permissions and terminal decisions are not left solely to the agent.
The headline is an operations warning
The database disagreement is not evidence that every agent task will fail, or that an agent’s completion message has no value. It is a warning against accepting the message as the final authority where external state is what the task is meant to change.
That changes how to read an agent demonstration. A list of tool calls and a polished final response are not sufficient evidence of task completion. The useful additional question is whether the workflow has a separately enforceable completion condition tied to the state the user actually needed.
Hugging Face’s title puts the problem in a memorable form: the agent said it was done; the database disagreed. The durable lesson is not to make the model argue harder. It is to ensure that the system which decides a job is finished can check something more substantial than the model’s own assertion.
- ai agents
- completion control
- agent reliability
- apple
- full-disk access
Questions
Does an agent saying it is done verify a database change?
No. Hugging Face’s post title explicitly contrasts an agent’s assertion that it was done with database state that disagreed. A completion message can be useful output, but it is not the same thing as evidence that the intended state exists.
What is completion control in an AI agent?
Completion control is the ability to decide whether an agent may end a task independently of the model that proposes actions. A preprint on arXiv reports that coupling this decision to a single LLM policy can allow unsupported completion claims to terminate execution.
Does the published material explain what caused the database mismatch?
No. The supplied Hugging Face material gives the post title but no account of the model, tool call, database operation or underlying cause. It supports the operational problem, not a diagnosis of that incident.
How do Apple’s permissions changes relate to completion control?
Apple’s permissions changes concern what an AI agent may access, while completion control concerns when its workflow may stop. Ars Technica — AI reports that Apple changed full-disk access permissions to curb abuse from AI agents, illustrating a separate control boundary.
Sources
Every page this piece was written from.
- The Agent Said It Was Done. The Database Disagreed. · huggingface.co
- DeReAct: Decomposed Reasoning and Acting for Reliable AI Agents · arxiv.org
- Apple changes full-disk access permissions to curb abuse from AI agents · arstechnica.com
About the author
odnoga Team
The odnoga team writes about artificial intelligence for the people who build with it: what shipped, what the research actually found, and what it means for the week ahead. Every piece names its sources.
More from this section
- Why cheaper AI models do not settle the token bill · 24 September 2026
- Gemini 4 Argon is announced but access is restricted · 2 October 2026
- AMD’s World Labs deal would bring a world-model lab in-house · 30 September 2026
- Report says OpenAI agents scanned UN site over 16,000 times · 29 September 2026
Putting AI into your own app?
odnoga is the AI gateway for apps built on Lovable and Supabase: one API for every model, with usage, limits and billing per user.
Attribute every AI request to the user who made it, cap their spend and bill them for it, with one prompt pasted into Lovable.
AI billing for Lovable appsCall any model from your Edge Functions with the Supabase user attached: metered, budgeted and traced per user, without a package to install.
The AI gateway for Supabase