- Main
Automated Visualization for Structural Form-Finding using Orchestrated Multimodal Machine Learning Agents
Abstract
Effective collaboration is fundamental to successful structural design. While engineers rely on numerical models and technical diagrams, architects prefer conceptual visualizations conveying design intent. Traditional rendering is laborious, hindering the rapid iteration in conceptual design. Although recent text-to-image machine learning (ML) models offer a promising alternative, existing workflows often require manual intervention to correct unnatural artifacts. This paper addresses this gap by proposing a novel, agent-workflow hybrid framework that introduces autonomous self-correction to the visualization process. We introduce an architecture where multiple ML agents are orchestrated in a reflective loop to autonomously generate, evaluate, and iteratively refine visualizations. Using a CAD view and text description as input, these multimodal agents collaborate to generate images. If the result is flawed, the system autonomously uses the critique to guide a new generation attempt. This agentic approach reduces manual rework and consistently produces high-fidelity results, enabling a more efficient and creative design exploration.
Many UC-authored scholarly publications are freely available on this site because of the UC's open access policies. Let us know how this access is important for you.