Octrees as an Explicit 3D Language
Summary
OctLLM represents 3D geometry as explicit SparseOctree occupancy tokens and routes mesh tokens through independent trainable branches alongside a frozen vision-language backbone, setting a new state of the art among unified multimodal LLMs for image-to-3D and render-grounded captioning while preserving general language ability.
View Cached Full Text
Cached at: 10/05/26, 04:46 AM
Paper page - Octrees as an Explicit 3D Language
Source: https://huggingface.co/papers/2610.02388
Abstract
Existing3Dlargelanguagemodels(LLMs)compromiseontwofronts:theycompressshapesintolatentcodebookindicesorcoordinatetext,whichremovesspatialstructurefromwhatthemodelobserves,andtheyacquirethe3Dmodalitybyfine-tuningthebackbone,whichoverwritesitsgenerallanguageability.WepresentOctLLM,whichaddressesbothlimitations.Geometryentersasanexplicit3Dsequenceofoctreeoccupancytokens.However,fulloctreesequencesgrowrapidlywithdepth;OctLLMthereforerandomlyemptiespenultimate-levelnodesandomitsdescendantswhilepreservingshape,yieldingashortercoordinate-anddepth-anchoredSparseOctree(S-Octree)forposition-awaremask-modelinggenerationand3Dunderstanding.Ontheotherfront,existingmethodsintroduceanewmodalitywithfullfine-tuningorLoRA,butfullfine-tuningiscostly,LoRAlimits3Dcapacity,andbothmodifythelanguagepathway.OctLLMinsteadadds3Dcapacityinparametersseparatefromthepretrainedones:meshtokensareroutedthroughindependenttrainablebranchesinasubsetofblockswhiletextandimagetokensretainthefrozenvision-languagepathway,andthetwostreamsinteractthroughsharedself-attention.Ittrainsfarfewerparametersthanfullfine-tuning,yetsetsanewstateoftheartamongunifiedmultimodalLLMs,loweringimage-to-3DFIDby17.4%andraisingrender-groundedcaptioningby28.7pointsoverShapeLLM-Omni,whilematchingthebackboneongenerallanguagebenchmarks.
View arXiv pageView PDFProject pageGitHub2Add to collection
Get this paper in your agent:
hf papers read 2610\.02388
Don’t have the latest CLI?curl \-LsSf https://hf\.co/cli/install\.sh \| bash
Models citing this paper0
No model linking this paper
Cite arxiv.org/abs/2610.02388 in a model README.md to link it from this page.
Datasets citing this paper0
No dataset linking this paper
Cite arxiv.org/abs/2610.02388 in a dataset README.md to link it from this page.
Spaces citing this paper0
No Space linking this paper
Cite arxiv.org/abs/2610.02388 in a Space README.md to link it from this page.
Collections including this paper0
No Collection including this paper
Add this paper to acollectionto link it from this page.
Similar Articles
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
Imagine3D-LLM proposes teaching multimodal LLMs to imagine a compact 3D representation of a scene (decoded from learnable summary tokens as 3D Gaussian Splatting via photometric reconstruction loss) before answering multi-view spatial reasoning questions, outperforming prior 3D-aware MLLM approaches on multiple benchmarks.
LLM Neuroanatomy III - LLMs seem to think in geometry, not language
Researcher analyzes LLM internal representations across 8 languages and multiple models, finding that concept thinking occurs in geometric space in middle transformer layers independent of input language, supporting a universal deep structure hypothesis similar to Chomsky's theory rather than Sapir-Whorf linguistic relativism.
Native and Compact Structured Latents for 3D Generation
This paper introduces O-Voxel, a new sparse voxel representation for 3D generative modeling that efficiently handles complex topologies and appearance, and trains large-scale flow-matching models with 4B parameters to achieve state-of-the-art generation quality.
Euclid-Omni : A Unified Neuro-Symbolic Framework for Plane Geometry
Euclid-Omni is a neuro-symbolic framework integrating LLMs, VLMs, and a symbolic solver to address plane geometry problems from calculations to Olympiad-level proofs, using synthetic data generation for training.
SciOrch: Learning to Orchestrate Expert LLMs for Solving Frontier Multimodal Scientific Reasoning Tasks
SciOrch presents an 8B vision-language model trained with MCTS to coordinate multiple expert LLMs for multimodal scientific reasoning, achieving superior performance while reducing API costs.