Octrees as an Explicit 3D Language

Hugging Face Daily Papers Papers

Summary

OctLLM represents 3D geometry as explicit SparseOctree occupancy tokens and routes mesh tokens through independent trainable branches alongside a frozen vision-language backbone, setting a new state of the art among unified multimodal LLMs for image-to-3D and render-grounded captioning while preserving general language ability.

Existing 3D large language models (LLMs) compromise on two fronts: they compress shapes into latent codebook indices or coordinate text, which removes spatial structure from what the model observes, and they acquire the 3D modality by fine-tuning the backbone, which overwrites its general language ability. We present OctLLM, which addresses both limitations. Geometry enters as an explicit 3D sequence of octree occupancy tokens. However, full octree sequences grow rapidly with depth; OctLLM therefore randomly empties penultimate-level nodes and omits descendants while preserving shape, yielding a shorter coordinate- and depth-anchored Sparse Octree (S-Octree) for position-aware mask-modeling generation and 3D understanding. On the other front, existing methods introduce a new modality with full fine-tuning or LoRA, but full fine-tuning is costly, LoRA limits 3D capacity, and both modify the language pathway. OctLLM instead adds 3D capacity in parameters separate from the pretrained ones: mesh tokens are routed through independent trainable branches in a subset of blocks while text and image tokens retain the frozen vision-language pathway, and the two streams interact through shared self-attention. It trains far fewer parameters than full fine-tuning, yet sets a new state of the art among unified multimodal LLMs, lowering image-to-3D FID by 17.4% and raising render-grounded captioning by 28.7 points over ShapeLLM-Omni, while matching the backbone on general language benchmarks.
Original Article
View Cached Full Text

Cached at: 10/05/26, 04:46 AM

Paper page - Octrees as an Explicit 3D Language

Source: https://huggingface.co/papers/2610.02388

Abstract

Existing3Dlargelanguagemodels(LLMs)compromiseontwofronts:theycompressshapesintolatentcodebookindicesorcoordinatetext,whichremovesspatialstructurefromwhatthemodelobserves,andtheyacquirethe3Dmodalitybyfine-tuningthebackbone,whichoverwritesitsgenerallanguageability.WepresentOctLLM,whichaddressesbothlimitations.Geometryentersasanexplicit3Dsequenceofoctreeoccupancytokens.However,fulloctreesequencesgrowrapidlywithdepth;OctLLMthereforerandomlyemptiespenultimate-levelnodesandomitsdescendantswhilepreservingshape,yieldingashortercoordinate-anddepth-anchoredSparseOctree(S-Octree)forposition-awaremask-modelinggenerationand3Dunderstanding.Ontheotherfront,existingmethodsintroduceanewmodalitywithfullfine-tuningorLoRA,butfullfine-tuningiscostly,LoRAlimits3Dcapacity,andbothmodifythelanguagepathway.OctLLMinsteadadds3Dcapacityinparametersseparatefromthepretrainedones:meshtokensareroutedthroughindependenttrainablebranchesinasubsetofblockswhiletextandimagetokensretainthefrozenvision-languagepathway,andthetwostreamsinteractthroughsharedself-attention.Ittrainsfarfewerparametersthanfullfine-tuning,yetsetsanewstateoftheartamongunifiedmultimodalLLMs,loweringimage-to-3DFIDby17.4%andraisingrender-groundedcaptioningby28.7pointsoverShapeLLM-Omni,whilematchingthebackboneongenerallanguagebenchmarks.

View arXiv pageView PDFProject pageGitHub2Add to collection

Get this paper in your agent:

hf papers read 2610\.02388

Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash

Models citing this paper0

No model linking this paper

Cite arxiv.org/abs/2610.02388 in a model README.md to link it from this page.

Datasets citing this paper0

No dataset linking this paper

Cite arxiv.org/abs/2610.02388 in a dataset README.md to link it from this page.

Spaces citing this paper0

No Space linking this paper

Cite arxiv.org/abs/2610.02388 in a Space README.md to link it from this page.

Collections including this paper0

No Collection including this paper

Add this paper to acollectionto link it from this page.

Similar Articles

Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering

Hugging Face Daily Papers

Imagine3D-LLM proposes teaching multimodal LLMs to imagine a compact 3D representation of a scene (decoded from learnable summary tokens as 3D Gaussian Splatting via photometric reconstruction loss) before answering multi-view spatial reasoning questions, outperforming prior 3D-aware MLLM approaches on multiple benchmarks.

LLM Neuroanatomy III - LLMs seem to think in geometry, not language

Reddit r/LocalLLaMA

Researcher analyzes LLM internal representations across 8 languages and multiple models, finding that concept thinking occurs in geometric space in middle transformer layers independent of input language, supporting a universal deep structure hypothesis similar to Chomsky's theory rather than Sapir-Whorf linguistic relativism.

Native and Compact Structured Latents for 3D Generation

Papers with Code Trending

This paper introduces O-Voxel, a new sparse voxel representation for 3D generative modeling that efficiently handles complex topologies and appearance, and trains large-scale flow-matching models with 4B parameters to achieve state-of-the-art generation quality.