Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Summary
Function2Scene generates 3D indoor layouts from functional descriptions by parsing user needs and applying design constraints through an iterative refinement loop combining geometric analysis, LLM reasoning, and VLM assessment, outperforming baselines in satisfying functional requirements.
View Cached Full Text
Cached at: 06/01/26, 03:18 AM
Paper page - Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Source: https://huggingface.co/papers/2605.30819
Abstract
Function2Scene generates 3D indoor layouts from functional descriptions by parsing user needs and applying design constraints through an iterative refinement process combining geometric analysis, language modeling, and visual assessment.
Mosttext-driven 3D indoor scene synthesismethods generate rooms from object-centric prompts, asking what furniture should be placed rather than how the space is used. Yet in real interior design, a layout is judged by how well it supports its occupants, e.g., theiractivitiesand physical needs. We introduce Function2Scene, a framework for generating 3D indoor layouts fromfunctional specifications, i.e.,natural-language design briefsdescribing who will use a room and what they need to do there. Given such a specification, our system parsesoccupant personasandactivities, derives a customized set offunctional design constraintsfrom ataxonomy of 17 criteriaspanning spatial, ergonomic, activity, and environmental considerations, and uses these constraints to guide layout generation. Rather than relying on an LLM to directly produce a final scene, Function2Scene performs iterative evaluation and refinement through a tool-augmentedcheck-and-repair loop, combininggeometric measurements,LLM-based contextual reasoning, andVLM-based visual assessment. Experiments on 30 professionally written interior-design cases show that Function2Scene produces layouts that better satisfy functional requirements than recent LLM-based scene synthesis baselines, with our results preferred in 94.3% of pairwise comparisons. Our work reframes text-driven indoor scene synthesis from placing plausible objects to designing spaces that support human use.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.30819
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.30819 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.30819 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.30819 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
How are you extracting transaction tables from Indian bank statement PDFs? Looking for open-source/on-prem approaches
The author is working on a Credit Underwriting AI Agent and is seeking open-source on-premise approaches to extract structured transaction tables from diverse Indian bank statement PDFs, facing challenges with layout variations and needing reliable extraction methods.
LLM usage in Debian neither endorsed nor prohibited
Debian project is voting on a General Resolution regarding the use of large language models in contributions, with proposals ranging from prohibition to endorsement. The discussion period has been extended, and voting is scheduled for August 2026.
@freeCodeCamp: Getting structured data your app can actually trust can be tricky. In this tutorial, Vineeth explains how to design sch…
This tutorial from freeCodeCamp explains how to design schemas, validate outputs, and handle failures to reliably extract structured data from LLMs, covering techniques like constrained outputs, retry loops, and streaming.
@FeitengLi: LLM 玩的好多技术 Povey 在 Zipformer 里都探索过
Z.ai 推出 GLM-5.3-Flash,这是一个具有 1M 代币上下文窗口的多模态 AI 模型,参数规模为 320B-A18B,并以 MIT 许可证发布。
@simplifyinAI: An open source project turns any LLM into a 3D character that talks, moves, and reacts in the browser, no game engine, …
An open-source project enables turning any LLM into a 3D character that talks, moves, and reacts in the browser with memory of emotional states and relationships, fully customizable and free to use.