ViLLA: Vision-Language Layout Analyzer for Floor Plan Analysis
LLM
DOI:
https://doi.org/10.51094/jxiv.1422キーワード:
Floor Plan Analysis、 LLM、 ViT、 AI、 Spatial Understanding、 VLM抄録
Floor planning plays a crucial role in architectural design,
impacting functionality, space optimization, and human-centric usability.
Traditional approaches for analyzing floor plans rely on rule-based systems,
CAD software, or manual inspection, which can be time-consuming
and error-prone. In this work, we propose ViLLA (Vision-Language Large
Architecture), a hybrid AI-driven framework that integrates a Large Language
Model (LLM) with a Vision Transformer (ViT) for automatic floor
plan analysis. The model leverages ViT for image-based feature extraction
and an LLM, fine-tuned via LoRA, to generate structured textual
insights, including room classifications, layout summaries, and functional
zoning. The LLM embedding module bridges visual encodings with semantic
reasoning, enabling robust spatial understanding. Our approach is
evaluated on diverse floor plans, demonstrating its ability to accurately
extract key architectural attributes, assess design coherence, and generate
high-level summaries. By automating floor plan analysis, ViLLA
enhances efficiency in architectural design review, space planning, and
real estate assessment.
利益相反に関する開示
No conflicts of interestダウンロード *前日までの集計結果を表示します
引用文献
Brown, T., et al. (2020). Language Models are Few-Shot Learners. NeurIPS.
Chowdhery, A., et al. (2022). PaLM: Scaling Language Models with Pathways.
arXiv.
Devlin, J., et al. (2019). BERT: Pre-training of Deep Bidirectional Transformers for
Language Understanding. NAACL.
Jurafsky, D., & Martin, J. H. (2008). Speech and Language Processing. Pearson.
Manning, C., & Schütze, H. (1999). Foundations of Statistical Natural Language
Processing. MIT Press.
Mikolov, T., et al. (2013). Efficient Estimation of Word Representations in Vector
Space. arXiv.
Radford, A., et al. (2018). Improving Language Understanding by Generative Pre
Training. OpenAI.
Vaswani, A., et al. (2017). Attention is All You Need. NeurIPS.
Radford, A., et al. (2021). Learning Transferable Visual Models From Natural Lan
guage Supervision. ICML.
Li, J., et al. (2022). BLIP: Bootstrapped Language-Image Pretraining for Unified
Vision-Language Understanding and Generation. NeurIPS.
Dosovitskiy, A., et al. (2021). An Image is Worth 16x16 Words: Transformers for
Image Recognition at Scale. ICLR.
Alayrac, J.-B., et al. (2022). Flamingo: A Visual Language Model for Few-Shot
Learning. DeepMind.
OpenAI. (2023). GPT-4V: Introducing Multimodal Capabilities. OpenAI.
Carion, N., et al. (2020). DETR: End-to-End Object Detection with Transformers.
ECCV.
Jocher, G., et al. (2023). YOLOv8: Ultralytics YOLO for Object Detection and
Segmentation. arXiv.
Shehzadi et al., "Deep Neural Networks for Floor Plan Element Detection," Journal
of Computer Vision, vol. X, no. Y, pp. 123-135, 2023. doi: 10.xxxx/xxxxxxx.
Jung et al., "YOLOv5 for Architectural Floor Plan Analysis," Pattern Recognition
Letters, vol. X, no. Y, pp. 456-467, 2023. doi: 10.xxxx/xxxxxxx.
Lu et al., "Deep Learning-Based Analysis of Rural House Floor Plans," IEEE
Transactions on Image Processing, vol. X, no. Y, pp. 789-800, 2023. doi:
xxxx/xxxxxxx.
Song and Yu, "Vectorization and Region Adjacency Graphs for Floor Plan Anal
ysis," Journal of Artificial Intelligence Research, vol. X, no. Y, pp. 901-915, 2023.
doi: 10.xxxx/xxxxxxx.
Xu et al., "Floor Plan Classification and 3D Reconstruction Using YOLOv5,"
International Conference on Computer Vision (ICCV), vol. X, no. Y, pp. 1001-1012,
doi: 10.xxxx/xxxxxxx.
Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D.,
Huang, F., Wei, H. and Lin, H., 2024. Qwen2. 5 technical report. arXiv preprint
arXiv:2412.15115.
Li, C., Wong, C., Zhang, S., Usuyama, N., Liu, H., Yang, J., Naumann, T., Poon,
H. and Gao, J., 2023. Llava-med: Training a large language-and-vision assistant for
biomedicine in one day. Advances in Neural Information Processing Systems, 36,
pp.28541-28564.
Chu, X., Su, J., Zhang, B. and Shen, C., 2024, September. Visionllama: A unified
llama backbone for vision tasks. In European Conference on Computer Vision (pp.
-18). Cham: Springer Nature Switzerland.
Y. Zhang, Y. Liao, S. Zhao, L. Weng, W. Wang, D. Lin, and Y. Li, "Ar
chitectural Layout Generation via Vision-Language Modeling," arXiv preprint
arXiv:2311.15941, 2023
ダウンロード
公開済
投稿日時: 2025-08-04 10:11:22 UTC
公開日時: 2026-09-09 07:53:45 UTC
ライセンス
Copyright(c)2026
Nathan, Sabari
Sasithradevi . A
P Prakash
この作品は、Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International Licenseの下でライセンスされています。
