We spent nine months piloting an AI tool to automate design reviews and catch BOM errors before release. In controlled tests with standardized drawings and a narrow set of manufacturing processes, the tool flagged about 85% of known issues—missing dimensions, tolerance conflicts, manufacturability risks. Leadership loved the demos, and we got budget to roll it out enterprise-wide.
Then reality hit. Production drawings came from multiple sites with different CAD standards, inconsistent metadata, and variable lighting on scanned markups. Accuracy dropped to maybe 50%, and half the flags were false positives that just annoyed engineers. Worse, the tool had no idea when supplier constraints or tooling availability had changed since the data refresh, so it approved BOMs that couldn’t actually be built. We ended up with rework cycles and engineers bypassing the system entirely, keeping their shadow spreadsheets.
The real lesson wasn’t that the AI was bad—it was that we skipped the prerequisites. Our constraint definitions were scattered across emails and tribal knowledge. Our BOM semantics were a mess: engineering, manufacturing, and procurement all used different structures with the same part numbers. We never clarified who owned validation decisions at each handoff. The pilot succeeded because we controlled the variables; production exposed that our data and processes weren’t AI-ready. We’ve since stepped back to fix data governance and semantic alignment before trying again.