Tag
OctLLM represents 3D geometry as explicit SparseOctree occupancy tokens and routes mesh tokens through independent trainable branches alongside a frozen vision-language backbone, setting a new state of the art among unified multimodal LLMs for image-to-3D and render-grounded captioning while preserving general language ability.
This paper presents an open-source simulator for peaked quantum circuits using sparse state vector representations with vectorized operations and hardware acceleration, enabling efficient classical simulation of circuits with sharp output distributions.
FLUX3D introduces a framework for high-fidelity image-to-3D Gaussian Splatting generation by enhancing representation learning and cross-modal alignment with diffusion-aligned structured latents and a sparse-structure-aware diffusion transformer, achieving state-of-the-art results.