← Blog

The local model loaded. The simulated arm never moved.

Two local Qwen3.5 35B workflow sessions stopped before simulated Reach execution. Their failures matter, while reaching performance remains untested.

Timeline date

Covers September 20, 2026. Reconstructed later from available records.

Published · Updated

On September 20, we loaded a local Qwen3.5 35B model for a simulated SO-101 reach. Across two reserved sessions, the workflow never got through to a dispatched motion command.

That was a failure of the model/application workflow under the tested configuration. It matters even though there was no executed reach from which to judge the arm's task performance.

The two sessions stopped at different boundaries.

The first failure was ours

The first reservation succeeded, but its new allocation identity was absent from the runner-event schema. The run stopped before an inference request.

No model response, motion command or physics step existed to evaluate. We added the missing allocation value and checked the compatibility repair without another inference request.

The surrounding application had failed before the model had any opportunity to attempt the task.

The second failure reached the tool contract

The next session used Ollama 0.34.2 with cloud disabled, temperature 0, an 8K context and a 512-token output bound. Thinking was disabled and keep-alive was 0. These were the recorded trial settings, not a general recommendation for the model.

The model received an observation and emitted a request through the expected motion-tool route. One argument failed the configured contract. The relevant fragment was:

{"schema_version": "1"}

The tool schema required the numeric value:

{"schema_version": 1}

These are excerpts of the version argument, not complete motion requests. Roboty validated the returned arguments before dispatch and rejected the string value.

The result recorded one inference request, zero tool calls, zero control steps and zero simulated physics time. Here “tool calls” means requests admitted for dispatch to the simulator, not every tool request emitted by the model. A request was emitted. Zero motion calls were dispatched.

What failed, and what was untested

Question What these sessions established
Did the workflow get through to task execution? No. Both sessions stopped before motion, at different stages
Did the second session return valid motion arguments? No. The version argument failed the configured contract
Did a dispatched motion reach the target? Untested. No motion command was dispatched
Does this support a model ranking or general reliability estimate? No. Two development sessions cannot answer that

Both failures remain visible when assessing whether this workflow was usable. They stay outside the denominator for executed reaching attempts, because the simulated arm never received a command.

We stopped rather than loosen the frozen contract or spend the four unused session slots. The model was not promoted as the default or taken on to cube-to-tray. Loading it and reaching its tool interface had not produced one executable motion.

Who should supply a protocol version?

The version marker was application protocol metadata, not a motion target. In this trial, Roboty exposed the motion schema to the model and expected it to return that marker with the other arguments. The recorded evidence does not establish why this was the best production interface.

Looking back, there are at least three design choices. Strict rejection keeps accepted requests exactly within the declared contract. An explicit normalization rule could admit a narrowly defined alternate representation, but would need testing and a new contract. An adapter could instead own protocol metadata and ask the model only for the fields it needs to choose.

None of those alternatives retroactively changes the rejected request. Ollama's tool-calling interface gets a response to the application boundary. The unresolved design question is how much protocol bookkeeping the model should have to get right before that response can become an action.

Revision notes

  • : Clarified the simulated setting, emitted versus dispatched requests, and the significance of both workflow failures.
  • : Assigned a separate editorial timeline date to space the five stories in reading order; original evidence and publication dates are unchanged.

Related reading