プレプリント / バージョン1

ViLLA: Vision-Language Layout Analyzer for Floor Plan Analysis

LLM

##article.authors##

  • Nathan, Sabari Couger Inc
  • Sasithradevi . A VIT Chennai
  • P Prakash Anna University

DOI:

https://doi.org/10.51094/jxiv.1422

キーワード:

Floor Plan Analysis、 LLM、 ViT、 AI、 Spatial Understanding、 VLM

抄録

Floor planning plays a crucial role in architectural design,
impacting functionality, space optimization, and human-centric usability.
Traditional approaches for analyzing floor plans rely on rule-based systems,
CAD software, or manual inspection, which can be time-consuming
and error-prone. In this work, we propose ViLLA (Vision-Language Large
Architecture), a hybrid AI-driven framework that integrates a Large Language
Model (LLM) with a Vision Transformer (ViT) for automatic floor
plan analysis. The model leverages ViT for image-based feature extraction
and an LLM, fine-tuned via LoRA, to generate structured textual
insights, including room classifications, layout summaries, and functional
zoning. The LLM embedding module bridges visual encodings with semantic
reasoning, enabling robust spatial understanding. Our approach is
evaluated on diverse floor plans, demonstrating its ability to accurately
extract key architectural attributes, assess design coherence, and generate
high-level summaries. By automating floor plan analysis, ViLLA
enhances efficiency in architectural design review, space planning, and
real estate assessment.

利益相反に関する開示

No conflicts of interest

ダウンロード *前日までの集計結果を表示します

ダウンロード実績データは、公開の翌日以降に作成されます。

引用文献

Brown, T., et al. (2020). Language Models are Few-Shot Learners. NeurIPS.

Chowdhery, A., et al. (2022). PaLM: Scaling Language Models with Pathways.

arXiv.

Devlin, J., et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for

Language Understanding. NAACL.

Jurafsky, D., & Martin, J. H. (2008). Speech and Language Processing. Pearson.

Manning, C., & Schütze, H. (1999). Foundations of Statistical Natural Language

Processing. MIT Press.

Mikolov, T., et al. (2013). Efficient Estimation of Word Representations in Vector

Space. arXiv.

Radford, A., et al. (2018). Improving Language Understanding by Generative Pre

Training. OpenAI.

Vaswani, A., et al. (2017). Attention is All You Need. NeurIPS.

Radford, A., et al. (2021). Learning Transferable Visual Models From Natural Lan

guage Supervision. ICML.

Li, J., et al. (2022). BLIP: Bootstrapped Language-Image Pretraining for Unified

Vision-Language Understanding and Generation. NeurIPS.

Dosovitskiy, A., et al. (2021). An Image is Worth 16x16 Words: Transformers for

Image Recognition at Scale. ICLR.

Alayrac, J.-B., et al. (2022). Flamingo: A Visual Language Model for Few-Shot

Learning. DeepMind.

OpenAI. (2023). GPT-4V: Introducing Multimodal Capabilities. OpenAI.

Carion, N., et al. (2020). DETR: End-to-End Object Detection with Transformers.

ECCV.

Jocher, G., et al. (2023). YOLOv8: Ultralytics YOLO for Object Detection and

Segmentation. arXiv.

Shehzadi et al., "Deep Neural Networks for Floor Plan Element Detection," Journal

of Computer Vision, vol. X, no. Y, pp. 123-135, 2023. doi: 10.xxxx/xxxxxxx.

Jung et al., "YOLOv5 for Architectural Floor Plan Analysis," Pattern Recognition

Letters, vol. X, no. Y, pp. 456-467, 2023. doi: 10.xxxx/xxxxxxx.

Lu et al., "Deep Learning-Based Analysis of Rural House Floor Plans," IEEE

Transactions on Image Processing, vol. X, no. Y, pp. 789-800, 2023. doi:

xxxx/xxxxxxx.

Song and Yu, "Vectorization and Region Adjacency Graphs for Floor Plan Anal

ysis," Journal of Artificial Intelligence Research, vol. X, no. Y, pp. 901-915, 2023.

doi: 10.xxxx/xxxxxxx.

Xu et al., "Floor Plan Classification and 3D Reconstruction Using YOLOv5,"

International Conference on Computer Vision (ICCV), vol. X, no. Y, pp. 1001-1012,

doi: 10.xxxx/xxxxxxx.

Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D.,

Huang, F., Wei, H. and Lin, H., 2024. Qwen2. 5 technical report. arXiv preprint

arXiv:2412.15115.

Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon,

H. and Gao, J., 2023. Llava-med: Training a large language-and-vision assistant for

biomedicine in one day. Advances in Neural Information Processing Systems, 36,

pp.28541-28564.

Chu, X., Su, J., Zhang, B. and Shen, C., 2024, September. Visionllama: A unified

llama backbone for vision tasks. In European Conference on Computer Vision (pp.

-18). Cham: Springer Nature Switzerland.

Y. Zhang, Y. Liao, S. Zhao, L. Weng, W. Wang, D. Lin, and Y. Li, "Ar

chitectural Layout Generation via Vision-Language Modeling," arXiv preprint

arXiv:2311.15941, 2023

ダウンロード

公開済


投稿日時: 2025-08-04 10:11:22 UTC

公開日時: 2026-09-09 07:53:45 UTC
研究分野
情報科学