The Unstoppable Robot

How does a robot know when it’s done?

Tianyu Li @ Dexmate, Oct 1st 2026

While building a robot agent, I ran into a question that sounded easy: how does the robot know it is done? Imagine asking it to clean a house, room by room. It needs to recognize when the kitchen is clean and move to the next room. Otherwise, you get a very clean kitchen and a very unfinished house. I saw a smaller version of this in RoboLab (Yang et al., 2026), using the available pi0.5 checkpoint (Physical Intelligence et al., 2025): in the red-mug example below, the robot reaches the goal but keeps moving when automatic stopping is disabled.

A hand-written rule can provide that “done” signal, but each new task may need a new rule. A general completion checker also needs to keep up with the robot. If it takes several seconds to confirm that a robot has reached a door, the robot may keep moving and overshoot before the answer arrives. We need fast checks, repeated often enough to catch completion—and cheap enough that frequent calls do not turn into a large bill. That is why Jev caught my attention: TypeSafe’s text-only model promises much faster, cheaper decisions than a full vision-language model (VLM). Could a small VLM describe the scene and let Jev decide when to move on? This study asks: how accurately, quickly, and cheaply can we tell that a robot is done?

In case you’re wondering: what is Jev?

Jev is TypeSafe’s model for making structured decisions. You give it the current situation as text or JSON and a question with defined answers—say, “finished or unfinished?”—and it returns a choice with probabilities for the options. It also supports yes/no judgments and scores. TypeSafe calls this a “System One” model, designed for fast decisions that software can use directly. The version we tested takes text only, so it needs something else to describe the camera images. We called it through OpenRouter; its median judgment time in our benchmark was about 0.17–0.18 seconds. That speed was the attraction. Getting the visual description ready is another part of the bill.

In case you’re wondering: what is RoboLab?

RoboLab is NVIDIA’s simulation benchmark for testing robot manipulation policies. It provides virtual scenes, tasks such as stacking and putting objects in containers, and built-in success checks. That gave us camera observations to show the monitor and a simulator label to compare its answer against. We used its pi0.5 integration to generate the recordings, with success-triggered stopping disabled so we could see what happened after the goal was first reached. The monitor never receives the simulator’s success label—that stays on the grading side.

Reorient red mug task: Put the red mug upright so that the opening is facing upwards.
Current · 0.0 / 60 s Unfinished
Scene cameraWrist camera
Simulator: Finished Unfinished
0 s10 s20 s30 s40 s50 s60 s
The robot reaches the goal but keeps moving.

01Experiment setup

Jev only reads text, so it needs someone to look at the robot for it. I used a vision-language model (VLM) to describe the scene. But if a VLM is already looking, why not let it decide? I compared three routes: send its description to Jev, ask the same VLM to judge its own description in a separate call, or ask the VLM to decide directly. Every judge also gets the task and measured robot state.

I tried five VLMs across a range of sizes and costs: GPT-6 Astra, Gemini 3.8 Flash, DeepSeek V4.1 Flash, Qwen3.8-27B, and Qwen3.5-9B running locally. I wanted to see whether a smaller VLM plus Jev could hold its own against a larger model doing the whole job.

What exactly did the models get?

Each check uses two images: now and exactly one second earlier. Each image combines a scene view and a wrist view, so the VLM sees four camera views in total. The task instruction stays the same across all three routes.

The describing VLM gets the images and task. Jev and the second VLM call get its description, the task, and measured joint positions, joint velocities, and end-effector pose at both moments. The direct VLM gets the images, task, and robot state together. The two description routes share the same saved description. None of these judges sees the simulator’s success label or future observations.

The cloud models run through OpenRouter; Qwen3.5-9B runs locally on an RTX 5090. Local Qwen plus Jev still includes an online Jev call. We first compare the routes without a checklist, then repeat the judgments with a checklist generated from each run’s initial scene.

Images
NowNow: scene camera left and wrist camera right
One second earlierOne second earlier: scene camera left and wrist camera right
Each image: scene + wrist camera
Task description
Put the red mug upright so that the opening is facing upwards.
Measured robot state
Joint positions
Joint velocities
End-effector pose
Now + one second earlier
VLM → Jev
Images+ task
→
VLMDescribe
→
Description+ task + robot state
→
JevJudge text
→
Finished?Yes/No
VLM → VLM
Images+ task
→
VLMDescribe
→
Description+ task + robot state
→
Same VLMJudge text
→
Finished?Yes/No
VLM direct
Images + task + robot state
→
VLMJudge images
⟶
Finished?Yes/No
How images, task instructions, and robot state reach each judge.

For data, we ran pi0.5 on 15 RoboLab tasks ten times each, giving us 150 recordings. We let each run continue even after success, then selected 500 moments: 250 finished and 250 unfinished, labeled by the simulator. Each check sees the current moment and one second earlier, with both scene and wrist views. The question is always: is the task finished now? Most comparisons use all 500 examples; Astra uses a matched 50-example subset.

How were the examples selected?

The 15 tasks cover three difficulty levels, with five tasks per level and ten runs per task. We record at 15 Hz and let every run reach its time limit. A moment is labeled finished only while the simulator’s completion condition holds; an earlier success does not permanently make the run finished.

Every run contributes an unfinished example. The finished examples come from all 45 successful runs, with at most seven per run and at least two seconds between them. We favor clearer observations rather than sampling uniformly in time. The set is balanced overall, but six tasks have no finished examples, so we cannot measure completion sensitivity for those tasks.

Astra uses 50 examples: 25 finished and 25 unfinished, from 50 different runs across all 15 tasks. Comparisons with Astra use those same examples for the other models. Each check starts with a fresh context. We inspected the recordings and labels during quality checks, so this is not a never-seen holdout, and two snapshots cannot establish every contact or ordering condition.

I measured accuracy, time for the full check, and API cost per 1,000 checks, aiming for checks under two seconds. This is an offline test on recordings; the monitor has not earned the stop button yet.

How should I read the charts?

Samples. In the three results charts, Astra uses the balanced 50-case subset; other VLMs use 500 cases. Jev uses the corresponding descriptions and cases. Its separate simulator-facts row uses all 500 cases.

Runtime. Online calls use OpenRouter, including Jev. Local Qwen runs on an RTX 5090 (32 GB VRAM), Ryzen 9 9950X (16 cores / 32 threads), and Ubuntu 24.04 LTS.

Metrics and reference lines. Accuracy is balanced accuracy: average recall for finished and unfinished cases. A constant answer scores 50%. Time is the median for the complete check, including description and judgment where applicable. Failed attempts and retries count; model loading and initial checklist generation do not. Dashed reference lines mark 50% accuracy, the two-second target, or zero change. Changes are in percentage points (pp).

Cost. API dollars per 1,000 complete checks include confirmed charges for failed attempts and retries. Each route is priced independently, even when descriptions are shared. Local hardware, electricity, and initial checklist generation are excluded. Recorded Gemini charges include OpenRouter’s 50% discount (verified October 1, 2026); Jev charges are separate. Local Qwen inference has no API fee, but Qwen → Jev still calls the hosted judge.

Error bars. All estimates show nominal 95% trajectory-bootstrap intervals from 10,000 resamples of whole runs within each task. Changes use paired resamples of the same cases. Comparisons are exploratory, without adjustment for multiple comparisons. Zero-width intervals may appear as a single cap or be invisible; they do not imply certainty.

Simulator comparison. Each VLM is compared with its own description → VLM path, without a checklist. Jev is compared with Gemini → Jev (71.8%), selected after evaluation as the strongest 500-case description → Jev result and held fixed in resampling. Checklist changes use the same cases and descriptions; positive changes favor the checklist.

Latency detail. Medians hide slow checks. Counting failed attempts and retries, the 95th-percentile (P95) check took 20.25 s for Gemini direct, 10.92 s for Gemini → Jev, 19.16 s for local Qwen → Jev, and 130.1 s for Astra → Jev (50 cases; a few very long retries).

02Results and Findings

VLM → JevVLM → VLMVLM Direct
AccuracyTime Cost$ Cost
OpenAI GPT-6 AstraOnline
VLM → Jev
76.0%
10.28 s
$32.24
VLM → VLM
76.0%
13.60 s
$51.28
VLM Direct
86.0%
4.01 s
$22.54
Gemini 3.8 FlashOnline
VLM → Jev
71.8%
5.39 s
$3.24
VLM → VLM
79.4%
8.42 s
$5.29
VLM Direct
81.4%
5.50 s
$4.50
DeepSeek V4.1 FlashOnline
VLM → Jev
55.8%
4.12 s
$0.62
VLM → VLM
51.6%
5.49 s
$0.83
VLM Direct
51.2%
2.13 s
$0.22
Qwen3.8-27BOnline
VLM → Jev
62.2%
6.41 s
$1.31
VLM → VLM
65.6%
8.01 s
$1.41
VLM Direct
52.6%
2.08 s
$0.18
Qwen3.5-9BLocal
VLM → Jev
61.4%
9.34 s
$0.11
VLM → VLM
51.6%
9.45 s
$0.00
VLM Direct
50.0%
0.36 s
$0.00
0%50%100%
0251015
$0$25$50
Accuracy, speed, and API cost across the monitoring pipelines.

Gemini and GPT-6 Astra give the most accurate completion judgments. Their direct answers are slower and more expensive than those of the lower-tier models, but the cheaper alternatives remain much less reliable. The trade-off is clear: the models that recognize completion best cost more to keep checking.

Jev helps the lower-tier models improve on their direct answers, but it does not improve on direct Gemini or Astra. Separating description from judgment can help weaker models, although that benefit is not unique to Jev: Qwen3.8 does better judging its own descriptions. For the strongest models, going straight from images to a decision works best.

Direct answers are usually fastest. Describing the scene takes most of the time in the two-stage pipelines, so Jev’s fast judgment cannot remove the main delay. Even the fastest online direct answers exceed our two-second target at the median. Local Qwen takes just 0.36 seconds, but always answers “unfinished” on this benchmark. Speed alone does not make a useful monitor.

Separating description from judgment

Simulator factsObject poses, motion, measured contacts · now + one second earlier+ Task instruction + measured robot state
VLM or JevJudge the text
Finished?Yes/No

A completion check has two jobs: recover what is happening in the scene, then decide whether it satisfies the task. A poor score could come from either step. To investigate, we bypassed the visual description and gave each judge physical measurements extracted from the simulator, alongside the same task instruction and robot state.

These measurements describe object positions, bounding envelopes, orientation relative to gravity, motion, and measured contacts at the current moment and one second earlier. They contain no success label, task-specific success code, or future information. We supply no checklist. The judge still has to connect the measurements to the task.

simInfo → VLMsimInfo → Jev
With simulator factsChange from descriptions
OpenAI GPT-6 AstraOnline
Sim − own descriptions
98.0%
+22.0 pp
Gemini 3.8 FlashOnline
Sim − own descriptions
94.8%
+15.4 pp
DeepSeek V4.1 FlashOnline
Sim − own descriptions
50.6%
-1.0 pp
Qwen3.8-27BOnline
Sim − own descriptions
73.6%
+8.0 pp
Qwen3.5-9BLocal
Sim − own descriptions
50.0%
-1.6 pp
JevOnline
Sim − Gemini descriptions
68.2%
-3.6 pp
0%50%100%
-100+10+20+30
Accuracy with simulator facts, compared with visual descriptions.

The simulator facts bring Astra and Gemini close to perfect accuracy. When the physical state is explicit, they can usually decide whether it satisfies the task. This suggests that recovering the scene from images is a major bottleneck, while their task reasoning is relatively strong on this benchmark. Qwen3.8 also improves, but DeepSeek and local Qwen remain near the constant-answer baseline: clearer evidence is not enough for every judge.

Jev’s score is slightly lower than with Gemini’s descriptions, although the confidence interval spans zero, so we cannot establish a decline. One hypothesis is that the simulator’s positions, orientations, and contact measurements are harder for Jev to digest than task-focused language. Rewriting the same measurements into simpler language would help test that explanation.

The diagnostic still has limits. These simulator facts omit exact mesh geometry used by some containment checks and the earlier history required by the ordered task. They also change the format and richness of the evidence, so remaining errors cannot all be assigned to reasoning.

Could a large model tell the monitor what to check?

Once per trajectory · initial scene
Initial image + taskScene and wrist views
Gemini 3.8 FlashWrite the checking rule
Cached checklistObjects, conditions, evidence limits
At each check · now and one second earlier↓ Reuse checklist
Each judge also gets: task + measured robot state + cached checklist
VLM → Jev
Saved VLM description
Jev
Finished?Yes/No
VLM → VLM
Same saved description
Same VLM
Finished?Yes/No
VLM Direct
Current + previous images
VLM
Finished?Yes/No

Perhaps the monitor needed a clearer definition of “done.” We asked Gemini 3.8 Flash to inspect the task and initial scene once per trajectory and write a completion checklist for the monitor to reuse. The checklist identifies relevant objects, required relations and counts, necessary robot states, and limits on what the evidence can establish. It uses the task and pre-action image, without future frames, success labels, or simulator success code.

We then give the same cached checklist to each judge, keeping the cases, images, task, robot state, and descriptions fixed. The describing VLM never receives it. The chart below shows accuracy with the checklist and the change from adding it.

VLM → JevVLM → VLMVLM Direct
With Gemini’s checklistChange from no checklist
OpenAI GPT-6 AstraOnline
VLM → Jev
64.0%
-12.0 pp
VLM → VLM
76.0%
+0.0 pp
VLM Direct
84.0%
-2.0 pp
Gemini 3.8 FlashOnline
VLM → Jev
74.2%
+2.4 pp
VLM → VLM
78.8%
-0.6 pp
VLM Direct
80.4%
-1.0 pp
DeepSeek V4.1 FlashOnline
VLM → Jev
55.2%
-0.6 pp
VLM → VLM
53.2%
+1.6 pp
VLM Direct
51.2%
+0.0 pp
Qwen3.8-27BOnline
VLM → Jev
64.4%
+2.2 pp
VLM → VLM
65.4%
-0.2 pp
VLM Direct
53.0%
+0.4 pp
Qwen3.5-9BLocal
VLM → Jev
61.0%
-0.4 pp
VLM → VLM
52.4%
+0.8 pp
VLM Direct
50.2%
+0.2 pp
0%50%100%
-20-100+5
How an initial checklist changes completion accuracy.

The checklist produces no consistent improvement. Most changes are small, while Astra → Jev becomes less accurate. For the description routes, extra guidance can clarify what to verify, but cannot recover missing facts. Explaining individual errors would require inspecting the checklist and decision together. Guiding the describer to collect the right evidence remains a separate experiment.

Conclusion

For now, we still need a strong model to make the judgment directly. Jev does not push the stronger models further: judging their descriptions, it does worse than letting them judge the images themselves. It does help the smaller models, making better calls from their descriptions than they make on their own, though still well short of a strong model judging directly. Even so, the strong models are not yet fast or cheap enough to check constantly, so there is still a gap before a completion monitor is truly usable.

03Final Thoughts

A quick decision is still the goal

I still like Jev’s basic idea: make a useful decision quickly, without writing an essay at every check. The robot does not need a philosophical discussion about the mug. Jev’s own judgment is fast and cheap, but the Jev-based system we tested does not yet deliver the accuracy, speed, and cost we want together. Generating the visual description takes most of the time.

My next bet is to train a small VLM specifically for the judging job, so it can give a short answer directly from the task, images, and robot state. RoboReward already explores this direction, training 4B and 8B vision-language models to score robot task outcomes (Lee et al., 2026). The question for our monitor is how well a small model generalizes to new tasks and scenes while staying within our time and budget. I would also want more than “finished” or “unfinished.” Is the robot getting closer, getting stuck, or undoing its own work? Learned rewards and progress signals offer useful starting points (Ma et al., 2023; Liang et al., 2026). A signal like that could help us decide whether to keep going or change strategy. We did not evaluate these models here; they are promising next comparisons.

Write once, check often

Another idea is to have a strong coding agent build the stopping checks before the task starts (Liang et al., 2022; Surís et al., 2023). It could wire existing perception tools into the checking code: SAM 3 (Carion et al., 2025) to segment and track objects, or LocateAnything (Wang et al., 2026) to locate objects from text descriptions. The code could combine those outputs with robot state, then check object counts, gripper release, or whether a relation has held for several frames. Pay for the coding once, then reuse the monitor throughout the run. Perception still costs computation, but we could avoid generating a fresh written scene report at every check.

The harder part is covering what can happen during a long, complex task. An object gets occluded, a grasp slips, a completed step gets undone, or a recovery changes the order of later steps. The generated program needs to handle these possibilities and their combinations, including missing or ambiguous observations. Each individual rule may be simple, but anticipating and testing all the relevant cases is hard. A coding agent could help, but “just add another if statement” can become a very long afternoon.

When should the robot give up?

Calling an attempt a failure seems harder than recognizing completion. My Claude put it nicely: “Failed isn’t something you see, it is a decision to give up, made relative to a budget.” That captures the recoverable cases: a dropped mug may still be one good grasp away from success. “Unfinished” alone gives no reason to stop trying. Left to itself, the robot can always keep calm and carry on.

In practice, I would give the monitor a few concrete reasons to stop or ask for help: an exhausted time or retry budget, a broken object or unsafe situation, or an unproductive loop. Confirming a loop needs memory; one snapshot cannot tell us this is the fifth identical attempt. Related work covers learned failure explanations (Duan et al., 2024), recovery from execution history (Liu et al., 2023), and asking for help under uncertainty (Ren et al., 2023). And a safety check should be able to trigger intervention before we decide that the task has failed.

04Citation

If this post is useful for your work, you can cite it as:

@misc{li2026unstoppablerobot,
  author       = {Li, Tianyu},
  title        = {The Unstoppable Robot},
  year         = {2026},
  howpublished = {Blog post}
}

05References

  1. Jenai Xuning Yang et al. (2026). RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies.
  2. Physical Intelligence et al. (2025). π0.5: a Vision-Language-Action Model with Open-World Generalization.
  3. Tony Lee et al. (2026). RoboReward: General-Purpose Vision-Language Reward Models for Robotics.
  4. Anthony Liang et al. (2026). Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons.
  5. Yecheng Jason Ma et al. (2023). LIV: Language-Image Representations and Rewards for Robotic Control.
  6. Jacky Liang et al. (2022). Code as Policies: Language Model Programs for Embodied Control.
  7. Dídac Surís et al. (2023). ViperGPT: Visual Inference via Python Execution for Reasoning.
  8. Nicolas Carion et al. (2025). SAM 3: Segment Anything with Concepts.
  9. Shihao Wang et al. (2026). LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding.
  10. Jiafei Duan et al. (2024). AHA: A Vision-Language-Model for Detecting and Reasoning Over Failures in Robotic Manipulation.
  11. Zeyi Liu et al. (2023). REFLECT: Summarizing Robot Experiences for Failure Explanation and Correction.
  12. Allen Z. Ren et al. (2023). Robots That Ask For Help: Uncertainty Alignment for Large Language Model Planners.