Our approach
How we describe the evidence
A plan, a software check, a simulation and a physical robot test answer different questions. We describe the evidence we have, the conditions behind it and what it leaves unresolved.
- Source report
- A maintainer or author reports a feature, demonstration or evaluation. Read the source and its conditions.
- Documented task
- A task appears in a benchmark or guide. Inclusion alone does not establish successful performance.
- Hypothesis
- A labeled inference or proposed experiment, with no result yet.
- Confidence
- Confidence applies to the scoped statement. High confidence in a reported failure does not mean high task performance.
- Unknown counts
- A rounded percentage without counts stays that way. We do not reconstruct a success count or imply statistical certainty.
- Related configurations
- SO-100, bimanual and mobile results retain their own scope. They do not silently become single-arm SO-101 capabilities.
History and privacy
The Blog uses one post per timeline day to keep the story readable. Each post also shows when the underlying work happened and when it was actually published or updated. Corrections preserve the earlier account and explain what changed. The older research library keeps its own review history, feed and public export.
Only approved evidence is public. Raw device identifiers, calibration, household recordings, private notes and reviewer identities are excluded. A report shared privately is not permission to publish it.
What later physical evidence must establish
Predeclare conditions and preserve every attempted trial, failures, assistance, exclusions and interruptions. Assisted completion is not unassisted success. A failed physical test can still be useful evidence.
Beginner-tested status requires a genuinely new participant completing the defined beginner workflow, including setup, teaching/training, restart and evaluation, without continuous live coaching. Extensively helped formative participants do not qualify.
An accessibility claim needs a baseline-informed meaningful threshold declared before evaluation, a fair comparison with the best reasonable current baseline, and controls for practice and learning effects. A faster later attempt alone is insufficient.