A desktop-tidying robot built during the Adventure X 2026 hackathon. Centered on a Galaxea A1Z robotic arm, the system combines dual cameras, YOLOE, ArUco desktop localization, geometric grasping, and ACT imitation learning to identify and return everyday objects. A visual console, MCP natural-language control, and a Bluetooth ring make the robot easy to supervise and trigger.
A robotic arm that understands the desktop and returns objects to their places.
About the Project
This desktop-tidying robot was built by Damien He and me during the Adventure X 2026 hackathon, July 22–26. A Galaxea A1Z arm performs the actions, a fixed camera localizes objects on the desk, a wrist camera provides close-range observations, and either geometric grasping or an ACT policy returns each item.
Perception, coordinate transforms, grasp planning, and safety are separated into a locally verifiable pipeline.
01
Dual camerasGlobal tabletop + wrist close-up
→
02
YOLOEOpen-vocabulary masks and geometry
→
03
TransformsPixels → table mm → robot base
→
04
PlanningPose, IK, and bounded trajectory
→
05
A1Z motionExecute after safety confirmation
CORE 01 · SPATIAL CALIBRATION
Four tags define the whole tabletop frame
The system avoids extrapolating one tag’s error across the desk. A central reference tag establishes scale; all 16 corners then constrain the final homography.
16 corners80 mm scaleRMS error gate
TAG 01
TAG 02
TAG 03
TAG 04 · REF
Target
x / mm
y / mm
Reference origin
01
Detect the dictionary
Try 4×4, 5×5, 6×6, 7×7, and Original; keep the set with most detections.
02
Bootstrap scale
Map the most central reference tag to an ideal 80 × 80 mm square.
03
Correct observations
Use a Kabsch rigid fit to snap every observed tag back to an ideal square.
04
Refit the whole desk
Solve again from 4 × 4 corners so constraints span the entire workspace.
Coordinate pipeline
IMAGEpixel (u, v)
Homography→
TABLEmillimeter (x, y)
Similarity fit→
A1Z BASEmeter (X, Y, Z)
At least three non-collinear tag centers align the table frame with the A1Z base. Invalid scale or residual error disables geometric grasping.
0.0005–0.0015 m/mmDefault RMS ≤ 10 mm
CORE 02 · GRASP GEOMETRY
The mask says both “what” and “how to grasp”
A YOLOE mask is reduced to its centroid, major axis, and minor axis. These points are projected into metric table space before heading and jaw opening are calculated.
Local geometry
Major axis / heading
Minor axis / jaw span
Centroid / grasp center
Yaw −23.8°Width 42 mm
CORE 03 · SEMANTIC GUARDRAIL
The VLM interprets; YOLOE owns coordinates
“Return the bottle lying on the left.”
VLMResolve semantic ambiguity
+
YOLOE maskOutput motion geometry
Coarse cloud-vision pixels never drive the arm directly.
CORE 04 · MOTION SAFETY
Every motion passes a safety gate
01
Calibration validResolution, bounds, residual
✓
02
Target reachableIK and joint soft limits
✓
03
Trajectory boundedSegment ≤ 1.40 rad
✓
04
Operator confirmsworkspace_clear
✓
Total span ≤ 2.75 radAt most 2 segments
Any failed gate stops before a motion command is sent.
CORE 05 · KEEP TIDY
Inventory means “in the right place,” not just “visible”
LIVE
Cup At home
Charging dock 8.4 cm offset
Pencil case Not detected
Return queue01 1 item pending
Global nearest matching assigns identical objects to distinct home points. Missing and unregistered objects are reported separately.
SYSTEM TOPOLOGY
Perception, decisions, and motion converge on one local core
Two cameras provide complementary views while FastAPI coordinates perception, planning, and safety. MCP and the Bluetooth ring are task inputs that reuse the same guarded execution path.
LOCAL CORE
Sensing & input
FIXED CAMERA Fixed cameraInventory · ArUco · YOLOE
WRIST CAMERA Wrist cameraClose view · ACT · multi-view
BLE IMU RING Bluetooth ringSingle / double-tap trigger
Local control core
BACKEND FastAPI Orchestrator
Unified task state and hardware access
CAMERACALIBRATIONPERCEPTIONPLANNINGSAFETYTASK STATE
Interface & motion
MCP Agent interfaceNatural language, same safety path
A1Z + G1Z Arm and gripperTrajectory · grasp · return
LOCAL-FIRST COMPUTE
YOLOE inferenceTransformsIK solvingTrajectory planningRobot control
OPTIONAL CLOUD
The VLM resolves semantics only; it never outputs motion coordinates