EEBench, a simulation-based evaluation framework, tested whether AI models can design functional circuits, verified by SPICE simulation and design checks against real-world tolerances, according to the project's published results. Models work in atopile, a declarative code format for components and connections, rather than navigating a CAD tool's interface.
In results dated September 1, Claude Opus 5 scored 61.6%, followed by Grok 4.6 at 57.1%, Claude Fable 5.1 at 56.4%, Claude Fable 5 at 54.3% and Claude Opus 4.8 Max at 51.4%. GPT-5.5 scored 42.3% and GPT-5.6 Sol scored 39.4%. No GPT-6 Astra results were available at publication, EEBench said.
The models "know much more about electronics than their output in conventional design tools tends to show," EEBench found, but real design involves trade-offs the models routinely miss. One submitted design used only 11.4 microfarads of effective capacitance at operating voltage when 545 microfarads were required, and it failed simulation despite building successfully, according to EEBench.
A benchmark score in the 50s or 60s reads as promising until you remember a failed circuit board is a wasted manufacturing run, not a retry. Hardware doesn't get software's cheap iteration loop, so a model that is right most of the time is not yet a model you can trust unsupervised.