I thought this was going to be a post about learning robotics.
It seems more like a post about hooking up an LLM to a pre-made robotic arm.
While that's interesting, I wouldn't have labeled the post "my robotics crash course."
Unfortunately, I think this is in some sense another example of LLMs substituting for actual learning. While I'm sure the author is learning something about robotics from doing these experiments, I doubt it's as much as he would have gotten from say reading a couple chapters of an introductory robotics for dummies book.
systemerror 33 minutes ago [-]
I think my approach to learning most things I'm interested in is to set a basic goal and learn exactly enough info to get me to that goal. If the goal is too ambitious, I'll set a more reasonable goal and try again. If I'm able to achieve the goal, set a more ambitious goal and learn more during that process. Trying to learn everything about robotics is not a reasonable goal at this time so I'm doing my best to set myself up for success.
jvanderbot 55 minutes ago [-]
True in the sense of learning fundamentals from the bottom up.
Only partially true in the sense of learning how the system might work and doing problem-directed learning.
Eventually, all the problems you'll encounter getting an arm automated are self-discovered, and if you stick with it, you'll learn all you need eventually.
So this is basically step 0.
41 minutes ago [-]
jvanderbot 5 hours ago [-]
First law of calibration: Make sure you do an error analysis.
Finding the table surface is pretty useless using a top-down view, even with April tags, because the range error to an april tag is much more than the bearing (pixel) error to the april tag. You basically have trouble observing the thing you're trying to measure.
If you do this again, ask your agent to conduct this analysis and make sure your desired calibration variables are observable with small error. A second camera from a 45deg angle or even on the table would go a lot further, but then of course other things become unobservable.
Nice workaround using a proxy for force sensing to get touch info, however. And neat project overall!
systemerror 4 hours ago [-]
This is very useful info! I've added the additional 45 degree camera and so far it's seems like there is a regression on skill but I think it's because it wasn't available in the earlier parts of the training.
Can you elaborate on the calibration variables comment? This sounds useful but how would I apply these observations?
jvanderbot 4 hours ago [-]
This is long and rambly - let's email (see profile) if you want to work through it together.
You're effectively trying to understand how motor inputs change the end hand position, and in particular, you want to know where the table top is so you can position the hand close to it to pick up/ put down.
This means you have some tuning to "learn" before you can apply a control policy / algorithm - and you should be careful how you phrase this so claude/ai can pick up the right vocabulary and bias towards good solutions.
Adding multiple views helps as follows:
0. Measure from multiple views the table top - Keep cameras stead and rigid, and ask claude to use opencv to do multi-view registration so the plane of the table is known precisely. Keep the cameras steady throughout this process - if they wiggle, you can do multi-view registration each measurement...
Paint the "finger tips" bright orange. Not kidding. Use a very flat chess board under the arm for your "workspace". Also not kidding.
1. Move arm to known position, the multiple cameras will measure the april tags movement AND THE FINGER TIP LOCATIONS. The chess board provides a very nice texture. Or a big flat texture of any kind helps here. Since we know the cameras and table positions, we're getting closer to knowing how the arm movements move the hand w.r.t. the table.
2. If your arm has encoders, then manually touch the table at several points, and claude will record the april tags + joint angles + finger locations.
2b. If your arm does not have encoders, then manually touch the table at several points, and record the april tags + finger locations only, but SPECIFY that the arm is now touching the table. --> Claude can use more opencv code to actually locate the touch point on the table. The measurement is noisy, but you will use many movements to figure it out.
Repeat many times, 10-20, multiple touch points. Then, ask claude to do an error analysis and suggest more touch points. Tell it to use system identification techniques / camera/arm calibration techniques. tell it to research these techniques and report errors.
When you're done, you have to do repeatability experiments - the key is this is now automatic. Claude picks joint angles for the arm, the fingers move, the april tags move, the multiple camers calibrate and record positions, and claude repeats. The manual steps are just the bootstrap - you should be automated now.
What you want is claude to output and ORDF of the whole system - bang you can now control the arm using off the shelf software, which claude is happy to set up for you.
When it comes time to pick things up, the multiple views will locate the object precisely and a control network or algorithm can plug in to control it via the ORDF.
jacquesm 2 hours ago [-]
What a fantastic comment. The information density alone is impressive, the way in which you took such a complicated process and laid it out in a few paragraphs makes it more so.
pj_mukh 5 hours ago [-]
Shoutout to people who don't overthink blog setups. "thisismypersonalblog.com" AND GO. It'll do the job.
On a technical note, your setup is not far from where professional setups are headed [1]. Astra is really doing an end-run around (for now) research Robotics setups.
I would rate that as an AI-assisted setup for a traditional research robotic workspace. No different than "Claude set up ci/cd for this project", just much more hands-on (pun intended). Software ate the world so anything that can generate software is going to make world-space interactions easier.
bambataa 4 hours ago [-]
I also got a SO-101 and mucked about a bit. I stopped partly because I realised my idea of putting a laser on the end of it had potential risks but also when you’re just prompting an agent to do things it’s hard to feel that you’re really “learning robotics”.
This post has given me some inspiration though!
ainch 4 hours ago [-]
Curious to see how they get on with VLAs. They sound great till you have to sit and record hundreds of teleop demos to teach them how to solve your task...
Mystery-Machine 2 hours ago [-]
Anyone tried using Jev for fast AI robot decisions? Would it make sense?
thomasikzelf 2 hours ago [-]
What kind of decisions though? jev does not output joint angles, and if it did it also needs to know how. It also does not take in images.
Rendered at 20:10:50 GMT+0000 (Coordinated Universal Time) with Vercel.
It seems more like a post about hooking up an LLM to a pre-made robotic arm.
While that's interesting, I wouldn't have labeled the post "my robotics crash course."
Unfortunately, I think this is in some sense another example of LLMs substituting for actual learning. While I'm sure the author is learning something about robotics from doing these experiments, I doubt it's as much as he would have gotten from say reading a couple chapters of an introductory robotics for dummies book.
Only partially true in the sense of learning how the system might work and doing problem-directed learning.
Eventually, all the problems you'll encounter getting an arm automated are self-discovered, and if you stick with it, you'll learn all you need eventually.
So this is basically step 0.
Finding the table surface is pretty useless using a top-down view, even with April tags, because the range error to an april tag is much more than the bearing (pixel) error to the april tag. You basically have trouble observing the thing you're trying to measure.
If you do this again, ask your agent to conduct this analysis and make sure your desired calibration variables are observable with small error. A second camera from a 45deg angle or even on the table would go a lot further, but then of course other things become unobservable.
Nice workaround using a proxy for force sensing to get touch info, however. And neat project overall!
Can you elaborate on the calibration variables comment? This sounds useful but how would I apply these observations?
You're effectively trying to understand how motor inputs change the end hand position, and in particular, you want to know where the table top is so you can position the hand close to it to pick up/ put down.
This means you have some tuning to "learn" before you can apply a control policy / algorithm - and you should be careful how you phrase this so claude/ai can pick up the right vocabulary and bias towards good solutions.
Adding multiple views helps as follows:
0. Measure from multiple views the table top - Keep cameras stead and rigid, and ask claude to use opencv to do multi-view registration so the plane of the table is known precisely. Keep the cameras steady throughout this process - if they wiggle, you can do multi-view registration each measurement...
Paint the "finger tips" bright orange. Not kidding. Use a very flat chess board under the arm for your "workspace". Also not kidding.
1. Move arm to known position, the multiple cameras will measure the april tags movement AND THE FINGER TIP LOCATIONS. The chess board provides a very nice texture. Or a big flat texture of any kind helps here. Since we know the cameras and table positions, we're getting closer to knowing how the arm movements move the hand w.r.t. the table.
2. If your arm has encoders, then manually touch the table at several points, and claude will record the april tags + joint angles + finger locations.
2b. If your arm does not have encoders, then manually touch the table at several points, and record the april tags + finger locations only, but SPECIFY that the arm is now touching the table. --> Claude can use more opencv code to actually locate the touch point on the table. The measurement is noisy, but you will use many movements to figure it out.
Repeat many times, 10-20, multiple touch points. Then, ask claude to do an error analysis and suggest more touch points. Tell it to use system identification techniques / camera/arm calibration techniques. tell it to research these techniques and report errors.
When you're done, you have to do repeatability experiments - the key is this is now automatic. Claude picks joint angles for the arm, the fingers move, the april tags move, the multiple camers calibrate and record positions, and claude repeats. The manual steps are just the bootstrap - you should be automated now.
What you want is claude to output and ORDF of the whole system - bang you can now control the arm using off the shelf software, which claude is happy to set up for you.
When it comes time to pick things up, the multiple views will locate the object precisely and a control network or algorithm can plug in to control it via the ORDF.
On a technical note, your setup is not far from where professional setups are headed [1]. Astra is really doing an end-run around (for now) research Robotics setups.
[1]: https://x.com/ihorbeaver/status/2104646736652447854?s=20
This post has given me some inspiration though!