A Comparative Analysis of Generative AI for Architectural Visualization: A Quantitative Comparison Using Objective Evaluation Metrics and the Architectural AI Score (AIS)


Creative Commons License

Cantemir V., Kurtuluş M., Cantemir E.

APPLIED SCIENCES, cilt.16, sa.19, ss.1-21, 2026 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 16 Sayı: 19
  • Basım Tarihi: 2026
  • Doi Numarası: 10.3390/app16199877
  • Dergi Adı: APPLIED SCIENCES
  • Derginin Tarandığı İndeksler: Applied Science & Technology Source, Scopus, Science Citation Index Expanded (SCI-EXPANDED), Compendex, INSPEC, Directory of Open Access Journals
  • Sayfa Sayıları: ss.1-21
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • İstanbul Gelişim Üniversitesi Adresli: Evet

Özet

This study presents a quantitative metric-based comparison of four state-of-the-art text-to-image generative AI systems for architectural visualization: OpenAI Image Generation, Imagen 4, Midjourney v8.1, and Stable Image Ultra. The models were evaluated using the Architectural Prompt Dataset (APD-50), a controlled, expert-informed prompt set consisting of 50 architectural prompts across five typologies: Modern Residential Buildings, Public Buildings, Religious and Cultural Architecture, Interior Design, and Facade and Sustainable Architecture. For each prompt, four images were generated per model–prompt combination, resulting in a corpus of 800 images. The generated images were analyzed using reference-free and image-level computer vision metrics, including CLIP Score for prompt alignment, Raw BRISQUE and Quality_BRISQUE for reference-free perceptual quality, Shannon Entropy as a diagnostic descriptor of visual information density, LPIPS as intra-prompt perceptual variation, and SSIM as intra-prompt structural similarity. In response to the limitations of treating entropy as an inherently positive indicator, Shannon Entropy was excluded from the revised composite score and retained only as a diagnostic metric. Alongside the individual metric results, this study proposes the Revised Architectural AI Score (AIS-R) as a first-stage, exploratory composite index for summarizing selected visual-output characteristics under controlled prompt conditions. AIS-R combines normalized CLIP, Quality_BRISQUE, and LPIPS components using an explicit and reproducible weighting structure. The purpose of AIS-R is not to replace professional architectural judgment or to establish a fully validated architectural-quality framework at this stage. Rather, it provides an initial operational formulation for comparing prompt alignment, reference-free perceptual quality, and intra-prompt perceptual variation within the limited scope of the APD-50 experimental setup. Linear mixed-effects models with prompt-level random intercepts were used as the primary inferential framework. Under the revised metric configuration, Midjourney v8.1 achieved the highest AIS-R score (0.6504 ± 0.0586), followed by Stable Image Ultra (0.6267 ± 0.0582), Imagen 4 (0.6069 ± 0.0707), and OpenAI Image Generation (0.5829 ± 0.0555). However, these differences should be interpreted as metric-based differences in selected visual-output characteristics rather than as evidence of architectural quality, constructability, spatial validity, scale accuracy, or professional workflow suitability. The study therefore represents an initial metric-formulation and controlled benchmarking stage, while future research should examine the correspondence between AIS-R and professional architectural judgments through expert evaluation, user studies, or workflow-based validation.