Today I tested the same problem a crane operator has: something heavy, something holding it up, and a decision about how the two meet. At four tonnes that decision has a standard attached. Take a few zeros off and the problem becomes easy to replicate in my living room.
I simplified an engineering problem to a 4 kg kettlebell, a rolling pin across two dining chairs, garden twine and one knot. I handed it to ChatGPT, Gemini and Claude. And then I actually tried building whatever my AI friends asked me to, following instructions to the tee (when possible).
The Experiment
Same prompt, same photo, same link, all three models. Verbatim:
“I have a 4 kg kettlebell that I want to suspend with a twine on this horizontal rolling pin using exactly 1 knot. Guide me on exactly how I can accomplish this task.
Twine: https://amzn.eu/d/0bwk6sO5
A photo of the kettlebell and rolling pin is attached.”
(The photo shows the rolling pin balanced across two dining chairs, the kettlebell on the rug underneath)
I ran once per model, deliberately. Because this isn't a test of accuracy, where you'd need repeat runs to see where a model lands. The finding is the direction each model took with an identical prompt written in natural human language, and that's visible in a single run.
And to be clear about what's being measured: this is a path each model chose when handed the apparatus and the goal and a prompt in natural human language, not what it's capable of if pushed. But the point is what it chose when it was only supplied with a goal, a constraint and the available apparatus.
The Results
I scored the answers on four parameters:
Information Gathering: Did it actually read the link and the photo?
Solution Elegance: Is the rig it designs mechanically sound: stable, balanced, secured?
Instruction Clarity: Could I understand what I was being asked to do?
Execution practicality: Could I actually do it, with two hands, in this room?

Observations
Both usable answers were easy enough to follow. Only one of them produced a rig that hangs straight.
Twine held both times. Two strands of 1 mm 2-ply jute carrying 4 kg off a rolling pin balanced on two dining chairs, no visible strain.
Claude's output needed me to solve a problem it created. Its step 3 is - ‘pull until the bell just lifts off the rug’; step 4 is ‘tie one knot.’ Holding 4 kg steady with one hand while tying a two-handed knot with the other is not a thing a person can do. I lowered the bell, threaded both ends, tied at leisure, and re-hung it. Obvious once you're standing there. Not in the instructions. And not cheating either, because nobody told the model I only had two hands — which is rather the point.
Claude's finished rig was worse. Nothing grips the handle, so the twine slides and the bell hangs crooked.
I also built out Claude and Gemini's versions and here are the pictures.
Claude's version

Gemini's version

The Analysis
Below is a detailed review of the results and findings:
Access isn't ability
All three read the photo. Only one demonstrably used the link. That's not three tiers of skill — vision runs inside the model, while fetching a web page is a request a website can refuse, and Amazon refuses.
Which leaves something genuinely useful: a model discussing your link is not evidence it opened your link. ChatGPT proved it opened the page with details only the page contains. Claude said outright that it was blocked. The third option — talking convincingly around a link you never opened — is the one you can't spot from the outside.
A blocked lookup didn't produce silence
Claude's 85 lb is a real number. It's the standard rating for 3-ply industrial jute, repeated across thousands of supplier pages. Mine is 1 mm 2-ply craft twine, which mostly isn't rated at all because it's not typically used for this application.
When the specific fact is missing, the sentence still has to finish, and what finishes it is whatever's most common (frequent) in the training text. The same mechanism explains why the puzzle was easy and the twine was hard. Loop topology is documented to death. How this cord behaves under load isn't written down by anyone, anywhere.
The differentiation was in the chosen mechanics and not the writing
Claude's answer looked richer at first glance with its sequencing, warning about tail length, a note that the knot eats 7–8 cm of loop, a flag that the pin can roll, and illustrations.
Gemini's contained a lot less, and actually included some jargon knot names which I thought I wouldn't be able to follow but was actually able to follow fairly easily!
Both got me to a hanging kettlebell without opening a second tab.
What set Gemini's rig apart was a single mechanical choice, made in one clause and never explained. Gemini cinches the twine onto the kettlebell handle with a girth hitch. Claude drapes it over. That's the whole difference, and that's what produced a balanced hanging load instead of the lopsided one Claude guided me to because a cinched loop can't slide along the handle, so the bell hangs level and stays level. A draped one migrates, and mine hangs crooked.
What's important is that this isn't a matter of taste. Riggers have names for both of these.
In rigging language - Gemini's is a choker hitch, Claude's is a basket. And the OSHA (US workplace safety regulator) standard operating rules say that loads in a basket hitch must be balanced to prevent slippage, and that slings must be securely attached to their loads. Mine wasn't, and it tilted.
There's a trade-off though - a choker is rated much lower for load carrying than a basket. The grippier option is the weaker one.
But in this case, for this load both would have worked with respect to load carrying capacity and Gemini's was the smarter choice for preventing slippage.
Claude had a practical gap I only found when I actually built it
The gap was physical rather than mechanical. Claude's step 3 is: “pull until the bell just lifts off the rug" and step 4 is: "tie one knot." Read that on a screen and it's fine. Try it and you need a third hand. My read on why is that most text (and therefore training data) records what was done, not so much the interim steps and back and forth of the effort it took to do it. This is also why the diagrams changed nothing, and why I split clarity from practicality in the results table. They are indeed two independent parameters.
All 3 models started with a constraint nobody set which limited their ability to be creative
This one bothered me most. One could cut three equal lengths of twine, braid them into a single cord, bring the two ends of the braid together, and tie one knot capturing all six tails at once. Nothing would unravel, because unravelling needs a free end to back out of and there isn't one left. Six strands would carry the bell instead of two. And there would still be exactly one knot. Entirely legal.
And this isn't hindsight talking. Hand me thin garden twine and a 4 kg weight and braiding is where I'd start, because I can see how thin the stuff is. That judgement doesn't require expertise. It requires being in the room with the material.
Claude got closest — it suggested coiling the twine into three laps before joining the ends. Same strand count, same single knot, and it's the only model that went looking for margin at all. But it still worked with what it had, not thinking about how the materials could be transformed before starting to build the rig.
Claude's idea was stronger but would fail all at once. The braid was slightly weaker but would less likely cause a spontaneous failure. And since in this example the bigger risk is abrasion and spontaneous failure the braid wins.
However the more important takeaway is that none of the models thought of it. My read is that the reason is that written instructions (which is what models learn from) almost always begin after somebody has already decided what the materials are, and there's very little text describing going through the brainstorming process of how to make what you have work. That's mostly a process that happens inside people's heads.
A few comparative side notes on the two solutions:

The Human Moat
Two of the three answers got a kettlebell into the air. One of them did it well. Each model did something better than the others.
That part I've written about before — which model gets which job, and being the person who makes that call is also relevant here.
But the more valuable human moat this experiment uncovered is critical evaluation. It means really going deeper into every solution to find the better of multiple viable solutions. And scrutinising every answer to see if there are any gaps.
It's harder than it sounds, and this experiment showed me why. The models weren't wrong. But one solution was clearly the better one practically. And looking at the answers in text I'd probably have guessed Claude's was the better solution. But as we now know it wasn't.
So how can you work better with a model on physical tasks?
There's quite a lot you can do actually. I am listing the three that would make the biggest difference here:
Describe the context and consequences: For most tasks you can't really convey what's at stake. For physical ones you can, and it's the context models most consistently lack - a photo of a living room didn't stop ChatGPT holding it to the safety standards of an industrial lift.
Give the model the information in a form it can actually read: A link is a request a website can refuse. Amazon refused, Claude filled the gap, and the gap got filled with a heavier twine than mine. Screenshot the spec instead.
Ask for the output type you learn from: Claude drew diagrams unasked; Gemini would have if I'd said so. Know how you take instructions in, and say it.
The Transcript
ChatGPT

Gemini

Claude


Sources and caveats
The experiment is worth repeating a few times if you want to check for variability in the outputs of a single model. That was not my goal here.
Sources: OSHA 1910.184 (Slings) for the choker and basket definitions and the balanced-load rule — the same standard ChatGPT cited at me.
