Why Text-to-Geometry AI Keeps Failing on Simple Parts

Generative AI trained on words and images, rather than physics, can’t understand the design intent compressed into something as “simple” as a lock washer.

Ask an AI design tool for a bracket and you will get something bracket-shaped. Ask it for a lock washer and watch what happens.

That should be the wrong way round. A bracket has more faces, more features and more ways to be wrong. A lock washer is a ring with a slit and a twist in it. A first-year student can draw one from memory. And yet the washer is the part that exposes the limits of every text-to-geometry system on the market, and understanding why is more useful than any benchmark score.

The Shape is Not the Part

A lock washer is not defined by its geometry. It is defined by what its geometry does under load. The split is not a notch; it is a spring. The free height above the seated ring is the preload the washer stores when the bolt is torqued. The cut ends are there to bite into the bearing face and the nut so the joint resists loosening under vibration. Change the free height by a few tenths of a millimeter and you have not made a slightly different washer; you have made a part that doesn’t do its job. Specify the wrong temper and the same geometry becomes an expensive shim.

READ MORE: What “Engineering-Grade AI” Actually Requires

None of that lives in the word “washer.”

Text-to-geometry systems learn a mapping from language to shape, because shape is the thing they can see. Every CAD file ever posted is a record of geometry. Almost none of them record intent. In the file, the dimension that mattered and the dimension somebody picked because it looked about right are indistinguishable. So a model trained on shape learns the visual signature of a lock washer, reproduces it faithfully, and hands you a part that looks correct and springs wrong.

Complex parts hide the failure. Simple parts do not.

This is why the demos are so convincing. Complex parts are forgiving in exactly the way simple parts are not. A generated bracket has a hundred dimensions and most of them genuinely do not matter, so an approximately right bracket is often a usable bracket. A generated lock washer has perhaps five dimensions and four of them are load-bearing in the literal sense. There is nowhere for the error to hide.

Standard parts are also where intent has been compressed hardest. Decades of testing and failure analysis are packed into a handful of numbers in a table, and the reasoning behind them was discarded long ago. When a system regenerates that part from a text description, it is not recalling a decision. It is guessing at an outcome whose derivation it never had.

The Feedback Loop is the Real Problem

Software people ran into a version of this and got lucky. Bad generated code throws an error. Bad generated geometry compiles fine. It renders, it exports, it passes a design review on a screen…and it fails on a test rig four months later, or in the field, or in a recall. Industry estimates have long put the cost of fixing a design error after release at 10-100 times the cost of catching it at the drawing. Generative tools push more decisions earlier while making them harder to inspect, which is a bad combination, and not one a bigger model fixes.

READ MORE: Leo AI: How CAD-Aware AI is Changing Mechanical Design and Engineering Workflows

What would actually fix it? Three things, none of them glamorous.

  1. Train on physics, not just pictures. A system that has seen the standards, the load cases and the failure modes behind a part class can reason about which dimensions are functional. One that has only seen finished solids cannot, however many it has seen.
  2. Make the tool cite itself. When a generated part comes back with the clause, the table or the internal drawing it came from attached, an engineer can check it in seconds. When it comes back as a naked solid, checking it costs about as much as drawing it, which erases the point of generating it.
  3. Reward the tool for refusing. This is the hard one for vendors. A system that says “I do not have enough to specify this correctly; here is what is missing,” is worth more in an engineering department than one that always produces something. Independent evaluations have put general-purpose chatbots wrong on roughly 46% of engineering questions, and the damage there is not the error rate, but the confidence. Engineers are trained to distrust a colleague who is never uncertain. The same standard should apply to software.

The Right Test

If you are evaluating any of this, do not ask for the impressive part—ask for the boring one. Ask for a lock washer for a specific bolt size, material and vibration environment, then ask the system why it chose that free height. The answer tells you what you are actually buying. If it gives you a number with a reason and a source, you have a tool that understands parts. If it gives you a number, you have a very good renderer. If it invents the reason, you have a problem.

The industry will get there. The parts that tell us we have arrived will not be the spectacular ones; they will be the ones a student could draw.

More content from Special Focus: CAD/CAM/CAE.

About the Author

Maor Farid

Co-founder and CEO, Leo AI

Dr. Maor Farid is the co-founder and CEO of Leo AI. He holds a PhD in mechanical engineering, was a Fulbright postdoctoral fellow at MIT and is a co-inventor on three U.S. patents in geometry-native AI.

Sign up for our eNewsletters
Get the latest news and updates

Voice Your Opinion!

To join the conversation, and become an exclusive member of Machine Design, create an account today!