Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Summary
Function2Scene generates 3D indoor layouts from functional descriptions by parsing user needs and applying design constraints through an iterative refinement loop combining geometric analysis, LLM reasoning, and VLM assessment, outperforming baselines in satisfying functional requirements.
View Cached Full Text
Cached at: 06/01/26, 03:18 AM
Paper page - Function2Scene: 3D Indoor Scene Layout from Functional Specifications
Source: https://huggingface.co/papers/2605.30819
Abstract
Function2Scene generates 3D indoor layouts from functional descriptions by parsing user needs and applying design constraints through an iterative refinement process combining geometric analysis, language modeling, and visual assessment.
Mosttext-driven 3D indoor scene synthesismethods generate rooms from object-centric prompts, asking what furniture should be placed rather than how the space is used. Yet in real interior design, a layout is judged by how well it supports its occupants, e.g., theiractivitiesand physical needs. We introduce Function2Scene, a framework for generating 3D indoor layouts fromfunctional specifications, i.e.,natural-language design briefsdescribing who will use a room and what they need to do there. Given such a specification, our system parsesoccupant personasandactivities, derives a customized set offunctional design constraintsfrom ataxonomy of 17 criteriaspanning spatial, ergonomic, activity, and environmental considerations, and uses these constraints to guide layout generation. Rather than relying on an LLM to directly produce a final scene, Function2Scene performs iterative evaluation and refinement through a tool-augmentedcheck-and-repair loop, combininggeometric measurements,LLM-based contextual reasoning, andVLM-based visual assessment. Experiments on 30 professionally written interior-design cases show that Function2Scene produces layouts that better satisfy functional requirements than recent LLM-based scene synthesis baselines, with our results preferred in 94.3% of pairwise comparisons. Our work reframes text-driven indoor scene synthesis from placing plausible objects to designing spaces that support human use.
View arXiv pageView PDFProject pageAdd to collection
Get this paper in your agent:
hf papers read 2605\.30819
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2605.30819 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2605.30819 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2605.30819 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
A dataset with 52 Text to image model evaluation [P]
A new benchmark dataset and evaluation methodology for 52 text-to-image models has been published, including results, a leaderboard, and a gallery to assess performance on challenging prompts.
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
Benchmarking shows that 4-bit quantized Qwen3.8 27B retains performance on benchmarks like Terminal-Bench 2.1 while fitting on 24GB GPUs, but 1-bit quantization causes severe degradation.
We recovered 575k crop labels from a decade of manual Photoshop work to automate book digitization - more data, ResNet-50, and higher resolution all failed; ten operator clicks per book beat them [P]
The author details a project on automating book digitization for Urdu literature by recovering crop labels from manual Photoshop work, finding that ten operator clicks per book outperformed scaling data or model complexity, and discusses retouching methods using neural nets and classical techniques.
IBM's new Granite 4.2 models ride the wave of interest in local LLMs
IBM has launched Granite 4.2, a family of open-weight large language models with variants up to 30B parameters, featuring a 128,000-token context window and agentic reinforcement learning for enhanced capabilities.
@MinLiBuilds: Really have to thank open-source contributors like Unsloth and SGLang. Model factories releasing weights is just the first step. These people tirelessly adapt, write kernels, reduce memory usage, improve speed, add tools, write tutorials, and even provide free compute resources. Without them, many so-called 'open-source models' are just...
This article praises open-source contributors like Unsloth and SGLang, who through optimizing tools and technologies enable ordinary people to fine-tune large language models such as Qwen3.8-27B on consumer-grade GPUs, lowering the barrier to AI development.