Finding the table surface is pretty useless using a top-down view, even with April tags, because the range error to an april tag is much more than the bearing (pixel) error to the april tag. You basically have trouble observing the thing you're trying to measure.
If you do this again, ask your agent to conduct this analysis and make sure your desired calibration variables are observable with small error. A second camera from a 45deg angle or even on the table would go a lot further, but then of course other things become unobservable.
Nice workaround using a proxy for force sensing to get touch info, however. And neat project overall!
Can you elaborate on the calibration variables comment? This sounds useful but how would I apply these observations?
You're effectively trying to understand how motor inputs change the end hand position, and in particular, you want to know where the table top is so you can position the hand close to it to pick up/ put down.
This means you have some tuning to "learn" before you can apply a control policy / algorithm - and you should be careful how you phrase this so claude/ai can pick up the right vocabulary and bias towards good solutions.
Adding multiple views helps as follows:
0. Measure from multiple views the table top - Keep cameras stead and rigid, and ask claude to use opencv to do multi-view registration so the plane of the table is known precisely. Keep the cameras steady throughout this process - if they wiggle, you can do multi-view registration each measurement...
Paint the "finger tips" bright orange. Not kidding. Use a very flat chess board under the arm for your "workspace". Also not kidding.
1. Move arm to known position, the multiple cameras will measure the april tags movement AND THE FINGER TIP LOCATIONS. The chess board provides a very nice texture. Or a big flat texture of any kind helps here. Since we know the cameras and table positions, we're getting closer to knowing how the arm movements move the hand w.r.t. the table.
2. If your arm has encoders, then manually touch the table at several points, and claude will record the april tags + joint angles + finger locations.
2b. If your arm does not have encoders, then manually touch the table at several points, and record the april tags + finger locations only, but SPECIFY that the arm is now touching the table. --> Claude can use more opencv code to actually locate the touch point on the table. The measurement is noisy, but you will use many movements to figure it out.
Repeat many times, 10-20, multiple touch points. Then, ask claude to do an error analysis and suggest more touch points. Tell it to use system identification techniques / camera/arm calibration techniques. tell it to research these techniques and report errors.
When you're done, you have to do repeatability experiments - the key is this is now automatic. Claude picks joint angles for the arm, the fingers move, the april tags move, the multiple camers calibrate and record positions, and claude repeats. The manual steps are just the bootstrap - you should be automated now.
What you want is claude to output and ORDF of the whole system - bang you can now control the arm using off the shelf software, which claude is happy to set up for you.
When it comes time to pick things up, the multiple views will locate the object precisely and a control network or algorithm can plug in to control it via the ORDF.
It seems more like a post about hooking up an LLM to a pre-made robotic arm.
While that's interesting, I wouldn't have labeled the post "my robotics crash course."
Unfortunately, I think this is in some sense another example of LLMs substituting for actual learning. While I'm sure the author is learning something about robotics from doing these experiments, I doubt it's as much as he would have gotten from say reading a couple chapters of an introductory robotics for dummies book.
Setting tiny goals, getting to them yourself, then setting a slightly larger goal, that's much more intense.
What I mean is that I'm not sure what "success" means in this context. There are already programs that control robot arms. I would think that success would mean that you managed to write a program that would control a robot arm.
You're an experienced programmer although maybe not a physics or mechanical person. Executing this would mean that you learned the mechanics (most people trying this don't have your programming experience, and have to try to do both things at once!)
Learning the mechanics would mean that you would have a good instinct to critique the AI when success in future projects would mean manifesting something you'd never seen before (and being happy you had AI to help you punch above your weight.) This would also give you a good foothold to understand more complex movements and coordination.
At least that's my perspective. The deliverable isn't on the table, it's in your brain. Asking the LLM to do it is like asking the LLM to copy a famous painting. You copy a famous painting so you can learn the movements of the person who painted it, not for the painting itself.
Only partially true in the sense of learning how the system might work and doing problem-directed learning.
Eventually, all the problems you'll encounter getting an arm automated are self-discovered, and if you stick with it, you'll learn all you need eventually.
So this is basically step 0.
On a technical note, your setup is not far from where professional setups are headed [1]. Astra is really doing an end-run around (for now) research Robotics setups.
[1]: https://x.com/ihorbeaver/status/2104646736652447854?s=20
This post has given me some inspiration though!