Projector Is All You Train
Summary
This paper demonstrates that training only the projector in multimodal large language models achieves strong performance on 3D tasks, avoids language model drift, and improves training efficiency compared to joint training methods.
View Cached Full Text
Cached at: 08/21/26, 10:09 AM
# Projector Is All You Train
Source: [https://arxiv.org/html/2608.19726](https://arxiv.org/html/2608.19726)
Nyx IskandarThanks:Equal Contribution\. Correspondence to:nyx@ramenvr\.comornyx@berkeley\.edu\.Affiliation:Ramen VREmail:[nyx@ramenvr\.com](mailto:)Saathvik Selvan11footnotemark:1Thanks:Work done while at Ramen VR\.Affiliation:University of California, BerkeleyEmail:[sselvan@berkeley\.edu](mailto:)
###### Abstract
The typical training process of a multimodal large language model \(MLLM\) involves adapting both the language model backbone and the projector between the backbone and a modality\-specific encoder\. We ask whether fine\-tuning the backbone of an MLLM is necessary to adapt it to a new modality\. Through experiments on 3D MLLMs, we find that training only the projector is sufficient to achieve strong multimodal performance relative to existing baseline models and our jointly trained MLLMs with the same encoder and backbone\. We also show that joint training leads to undesirable drift in existing capabilities of the language model, which projector\-only training avoids by definition\. Furthermore, projector\-only training has approximately twice the training sample throughput of joint training\. We validate our findings across different language model backbones via 3D classification and captioning benchmarks as well as standard benchmarks evaluating language, vision, and spatial reasoning capabilities\.
## 1Introduction
Multimodal large language models \(MLLMs\) extend pretrained language models with representations from additional modalities, such as images, audio, or 3D data\[[1](https://arxiv.org/html/2608.19726#bib.bib3),[4](https://arxiv.org/html/2608.19726#bib.bib4),[22](https://arxiv.org/html/2608.19726#bib.bib5),[32](https://arxiv.org/html/2608.19726#bib.bib13),[20](https://arxiv.org/html/2608.19726#bib.bib14)\]\. A common architecture visualized in Figure[1](https://arxiv.org/html/2608.19726#S1.F1)maps features from a pretrained modality\-specific encoder into the language model’s embedding space through a learned projector\. Training typically proceeds in two stages: the projector is first trained to align encoder features with the language model \(LM\), after which both the projector and LM are adapted on multimodal instruction data\[[43](https://arxiv.org/html/2608.19726#bib.bib2),[9](https://arxiv.org/html/2608.19726#bib.bib6),[13](https://arxiv.org/html/2608.19726#bib.bib7),[21](https://arxiv.org/html/2608.19726#bib.bib8),[22](https://arxiv.org/html/2608.19726#bib.bib5)\]\.
We question whether adapting the LM in this second stage is necessary\. For 3D MLLMs, we find that training the projector alone is sufficient to achieve strong multimodal performance\. Across multiple LM backbones, models trained with a frozen LM perform comparably to existing baseline models on 3D classification and captioning benchmarks\[[43](https://arxiv.org/html/2608.19726#bib.bib2),[44](https://arxiv.org/html/2608.19726#bib.bib1),[3](https://arxiv.org/html/2608.19726#bib.bib22),[37](https://arxiv.org/html/2608.19726#bib.bib24)\]\. Furthermore, we independently executed training runs where both the projector and LoRA adapters\[[19](https://arxiv.org/html/2608.19726#bib.bib11)\]of the backbone are jointly trained in a one\-stage process using the same dataset as the projector\-only training runs\. Comparing the 3D classification performance throughout training under matched compute budgets, we find that projector\-only training is consistently competitive with joint training when using the same point cloud encoder and LM backbone\. These results suggest that training the projector alone can be sufficient to adapt an MLLM to a new modality\.
In our experiments, projector\-only training is approximately twice as fast as joint training\. Moreover, joint training introduces regressions in LM performance across language, vision, and spatial reasoning benchmarks\. This drift is guaranteed to be absent in projector\-only training by construction, since the backbone weights are never updated\.
Figure 1:Architecture of 3D MLLMs in this paper\. This figure depicts projector\-only training, where all parameters are frozen except those of the projector\. An alternative training regime is joint training, where all parameters are frozen except those of the projector and LoRA adapters of the LLM backbone\.
## 2Related Work
3D MLLMs\.Existing 3D multimodal large language models \(MLLMs\) differ primarily in their point\-cloud representations, modality\-alignment strategies, and forms of instruction supervision\. ShapeLLM\[[32](https://arxiv.org/html/2608.19726#bib.bib13)\]targets interaction\-oriented 3D understanding by connecting a LLaMA backbone to ReCon\+\+, a point\-cloud encoder extended from ReCon\[[31](https://arxiv.org/html/2608.19726#bib.bib51)\]\. PointLLM\[[43](https://arxiv.org/html/2608.19726#bib.bib2)\]directly maps features from a Point\-BERT encoder pretrained under ULIP\-2\[[45](https://arxiv.org/html/2608.19726#bib.bib27),[46](https://arxiv.org/html/2608.19726#bib.bib28)\]through an MLP projector into a decoder\-only LLM\. Its two\-stage training first aligns point and language representations using 660k brief object captions and then jointly instruction\-tunes the projector and LLM using 70k GPT\-4\-generated complex interactions; the work also proposes generative 3D classification and captioning benchmarks with human\- and GPT\-based evaluation\. PointLLM\-V2\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\]expands the dataset to approximately 1\.7M point\-text samples covering 3D objects contained in Objaverse\-XL\[[10](https://arxiv.org/html/2608.19726#bib.bib21)\]in addition to Objaverse\[[11](https://arxiv.org/html/2608.19726#bib.bib15)\], and enriches the instruction space with coordinate\-based part\-referring questions\. PointLLM\-V2 also replaces the Vicuna 7B and 13B backbones in PointLLM with Llama\-3\.1\-8B\-Instruct\. PointLLM\-R\[[3](https://arxiv.org/html/2608.19726#bib.bib22)\]equips PointLLM with explicit chain\-of\-thought reasoning using PoCoTI, a 55k\-sample dataset constructed by first evaluating and refining point\-text instructions and then synthesizing geometrically grounded reasoning traces through Human\-in\-the\-Loop Prompt Optimization\. Finally, MiniGPT\-3D\[[37](https://arxiv.org/html/2608.19726#bib.bib24)\]emphasizes training efficiency by using a pretrained 2D vision\-language model \(VLM\) as an intermediate semantic bridge between point clouds and an LLM\. It combines a four\-stage cascaded alignment procedure with a mixture\-of\-query\-experts module, LoRA, and normalization\-layer fine\-tuning\. In this paper, we use the PointLLM family as our baseline for model architecture and dataset, as well as the evaluation strategy it pioneered which has since been adopted in 3D MLLM literature\.
Point\-BERT encoder\.The encoder used by the PointLLM family is Point\-BERT, pretrained under ULIP\-2\[[46](https://arxiv.org/html/2608.19726#bib.bib28),[45](https://arxiv.org/html/2608.19726#bib.bib27)\]\. It receives image\-text\-point contrastive supervision, aligning its representation to a language\-adjacent space\. Architecturally, Point\-BERT adapts the BERT\[[12](https://arxiv.org/html/2608.19726#bib.bib16)\]recipe to point clouds\. Input points are partitioned into 512 local patches by farthest\-point sampling followed by k\-nearest neighbor grouping\. After that, a mini PointNet\[[30](https://arxiv.org/html/2608.19726#bib.bib17)\]embeds each patch, and the resulting embeddings are passed through a multi\-head self\-attention transformer\. The output is a 513×\\times384 set of tokens \(512 patch tokens plus the class token\) which we pass to the projector\. Point\-BERT natively accepts each point as a 6\-dimensional vector, where the first three elements represent x\-, y\-, and z\-coordinates, and the last three elements represent the color of that point in RGB format\.
PointLLM dataset\.We use point clouds released with the PointLLM dataset, produced by surface\-sampling each mesh and transferring per\-point color from its texture data such that each object is represented as a colored point cloud ofn=8192n=8192points withd=6d=6channels, the first three channels being position and the last three RGB\[[43](https://arxiv.org/html/2608.19726#bib.bib2),[11](https://arxiv.org/html/2608.19726#bib.bib15)\]\. For supervision, we use annotations from the PointLLM\-V2 dataset, which upgrades the original’s text\-only supervision with a vision\-based pipeline for complex instruction generation\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\]\. This dataset is split into two subsets corresponding to the two\-stage training process\. The Stage 1 subset contains brief description prompts that ask for a one\-line caption of the object supplied by Cap3D\[[24](https://arxiv.org/html/2608.19726#bib.bib9)\]; the Stage 2 subset contains more complex prompts spanning long captioning, conversational question\-answering, and part\-referring\. Prompts were either drawn from a fixed collection or generated by GPT\-4o\[[27](https://arxiv.org/html/2608.19726#bib.bib32)\]\.
## 3Methodology
This section explains our MLLM architecture and two training regimes\. We also outline and motivate our experiment setup\.
### 3\.1Architecture Design
3D MLLMs are generative models that output a sequence of text tokens given an input point cloud and an input sequence of text tokens\. Architecturally, ours consist of a pretrained point cloud encoderfencf\_\{enc\}, a projectorfprojf\_\{proj\}, and a pretrained LM backboneflmf\_\{lm\}\. Onlyfprojf\_\{proj\}\(and, in joint training, low\-rank adapters insideflmf\_\{lm\}\) is ever optimized, whilefencf\_\{enc\}is frozen in every experiment\.
3D encoder\.The encoderfencf\_\{enc\}takes as input a point cloudP∈ℝn×dP\\in\\mathbb\{R\}^\{n\\times d\}and outputs a sequence of point featuresX∈ℝm×cX\\in\\mathbb\{R\}^\{m\\times c\}, wherennis the number of points,ddis the feature dimension of each point,mmis the number of point features, andccis the dimension of each point feature\.
Projection module\.The basic purpose of the projectorfprojf\_\{proj\}is to map point featuresXXfrom the output offencf\_\{enc\}into the embedding space offlmf\_\{lm\}\. This means thatfprojf\_\{proj\}takes as input point featuresXXand outputs point tokensY∈Rm×c′Y\\in R^\{m\\times c^\{\\prime\}\}, wherec′c^\{\\prime\}is the dimension of the token embeddings of the corresponding LM\. Following the MLP projector configuration recommended by PointLLM\-V2, we fixfprojf\_\{proj\}to a 3\-layer MLP with hidden dimensionsd1=1024d\_\{1\}=1024andd2=2048d\_\{2\}=2048and GELU\[[18](https://arxiv.org/html/2608.19726#bib.bib18)\]activations between consecutive linear layers\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\]\.
LM backbone\.The language modelflmf\_\{lm\}is a decoder\-only transformer\[[38](https://arxiv.org/html/2608.19726#bib.bib19)\]taking in a sequence of token embeddingsZ∈Rk×c′Z\\in R^\{k\\times c^\{\\prime\}\}, wherekkis the total number of tokens\.ZZcan contain token embeddings defined in the vocabularyVVof the language model as well as point tokensYYoutputted byfprojf\_\{proj\}\. WheneverZZincludesYY,YYis always surrounded by special tokens<\|pc3d\_start\|\>and<\|pc3d\_end\|\>, which are added toVVas input\-only tokens\. The output offlmf\_\{lm\}is a sequence of contextual hidden statesZ^=flm\(Z\)∈Rk×c′\\hat\{Z\}=f\_\{lm\}\(Z\)\\in R^\{k\\times c^\{\\prime\}\}\. Due to causal masking, each hidden statez^i\\hat\{z\}\_\{i\}depends only on previous embeddingsZ≤iZ\_\{\\leq i\}\. A linear vocabulary headfvocab:Rc′→R\|V\|f\_\{vocab\}:R^\{c^\{\\prime\}\}\\rightarrow R^\{\|V\|\}maps eachz^i\\hat\{z\}\_\{i\}to a logitli=fvocab\(z^i\)l\_\{i\}=f\_\{vocab\}\(\\hat\{z\}\_\{i\}\)\. Under greedy decoding, the predicted next token isz~i\+1=arg maxw∈Vsoftmax\(li\)\[w\]\\tilde\{z\}\_\{i\+1\}=\\text\{arg max\}\_\{w\\in V\}\\text\{softmax\}\(l\_\{i\}\)\[w\]\. In this paper, the three choices for the LM backbone are Qwen3\.5\-4B\[[33](https://arxiv.org/html/2608.19726#bib.bib25)\], Qwen3\.5\-9B\[[33](https://arxiv.org/html/2608.19726#bib.bib25)\], and Llama\-3\.1\-8B\-Instruct\[[23](https://arxiv.org/html/2608.19726#bib.bib26)\]\.
### 3\.2Training Regimes
We train the relevant parameters of our MLLMs by minimizing the negative log\-likelihood, or equivalently the token\-level cross\-entropy loss, over training token sequences, which is the causal language modeling objective\[[34](https://arxiv.org/html/2608.19726#bib.bib20)\]\. For a sequencex1:Tx\_\{1:T\}, the loss minimized to optimize parametersθ\\thetaisℒ\(θ\)=−1T∑logPθ\(xt\+1\|x≤t\)\\mathcal\{L\}\(\\theta\)=\-\\frac\{1\}\{T\}\\sum\{\\log\{P\_\{\\theta\}\(x\_\{t\+1\}\|x\_\{\\leq t\}\)\}\}\. We ran response\-only supervised fine\-tuning, masking prompt\-token labels such that only assistant\-response tokens directly contribute to the loss\.
Inprojector\-only training, only the parameters of the projector are optimized\. Injoint training, both the parameters of the projector and the LoRA adapters\[[19](https://arxiv.org/html/2608.19726#bib.bib11)\]of the LM backbone are optimized, where the adapters are attached to all linear layers of the backbone\. Appendix[A](https://arxiv.org/html/2608.19726#A1)contains more details on the hyperparameters\. The dataset used for both training regimes is the same subset of the mixed Stage 1 and Stage 2 data in PointLLM\-V2\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\]corresponding to objects in Objaverse\[[11](https://arxiv.org/html/2608.19726#bib.bib15)\]only\. In other words, our training runs do not distinguish between Stage 1 and Stage 2 data when sampling as we do not follow the two\-stage training paradigm this dataset was originally constructed for\.
### 3\.3Experiment Setup
We execute training runs on MLLMs with different backbones for both projector\-only and joint training as shown in Table[1](https://arxiv.org/html/2608.19726#S3.T1)\. Every MLLM is trained under an identical wall\-clock time budget of 16 hours on a single A100\-80GB GPU\. We match by wall\-clock time rather than by optimizer steps, since the latter would give joint training approximately2×2\\timesthe compute of similar projector\-only training runs\. This reflects a more realistic GPU budget a practitioner actually faces when deciding which method to use\. The 16 hours are measured from the first optimizer step, excluding dataset preprocessing and model loading time\. All runs use identical batch and sequence settings\. We report additional training details, including hyperparameters and throughput, in Appendix[A](https://arxiv.org/html/2608.19726#A1)\.
Table 1:Configurations of MLLM variants\. All variants are trained for the full 16 GPU hours\. We save checkpoints at 2, 4, 8, 12, and 16 hours\.Our experiments vary the language model backbone but not the encoder\. This is to control our experiments and make the trained MLLMs comparable to the baseline PointLLM family which all use the Point\-BERT encoder\. Furthermore, the encoder remains perpetually frozen under both training regimes; on the other hand, the backbones are ablated as they are updated in joint training\. By and large, we are more interested in validating that our findings hold for different language models\.
3D understanding evaluations\.We evaluate model checkpoints at 2\-, 4\-, 8\-, 12\-, and 16\-hour thresholds against benchmarks assessing 3D understanding\[[3](https://arxiv.org/html/2608.19726#bib.bib22),[43](https://arxiv.org/html/2608.19726#bib.bib2)\]\. These include generative zero\-shot 3D object classification benchmarks and a 3D object captioning benchmark first proposed by PointLLM\. The classification benchmarks use point clouds from ModelNet40\[[41](https://arxiv.org/html/2608.19726#bib.bib10)\], Objaverse\[[11](https://arxiv.org/html/2608.19726#bib.bib15)\]and, since PointLLM\-R, OmniObject3D\[[40](https://arxiv.org/html/2608.19726#bib.bib23)\]\. The captioning benchmark uses Objaverse point clouds\. For evaluations, the classification benchmarks use LLM\-as\-a\-Judge to score output generation correctness; in this paper, the judge used is GPT\-5\.6 Luna\[[28](https://arxiv.org/html/2608.19726#bib.bib29)\]and existing models are re\-run and re\-scored under said judge\. The Objaverse captioning benchmark also uses LLM\-as\-a\-Judge that is given multi\-view images of the 3D object to determine a correctness and hallucination score for the generated open\-ended caption, aggregated to a precision score as computed by PointLLM; in this paper, we use an ensemble of judges, namely GPT\-5\.6 Luna\[[28](https://arxiv.org/html/2608.19726#bib.bib29)\], Claude Haiku 4\.5\[[2](https://arxiv.org/html/2608.19726#bib.bib30)\], and Gemini 3\.5 Flash\-Lite\[[17](https://arxiv.org/html/2608.19726#bib.bib31)\]\. More details on the evaluation configuration, including the judge prompts, are provided in Appendix[B](https://arxiv.org/html/2608.19726#A2)\.
LM backbone evaluations\.To determine the extent to which existing capabilities of the LM drift after LoRA fine\-tuning for joint training variants, we load and merge the trained LoRA adapters into the base LM weights and run standard benchmarks against the LM backbone\. These include language benchmarks like MMLU\-Pro\[[39](https://arxiv.org/html/2608.19726#bib.bib33)\]and WinoGrande\[[36](https://arxiv.org/html/2608.19726#bib.bib36)\], vision benchmarks like MMMU\-Pro\[[48](https://arxiv.org/html/2608.19726#bib.bib37)\]and BabyVision\[[5](https://arxiv.org/html/2608.19726#bib.bib34)\], and spatial intelligence benchmarks like ERQA\[[16](https://arxiv.org/html/2608.19726#bib.bib35)\]and LingoQA\[[25](https://arxiv.org/html/2608.19726#bib.bib38)\]\. We also ran these benchmarks against the base LM itself to obtain a verified baseline; the base LM corresponds to the backbone of projector\-only\-trained MLLMs\. Since Llama\-3\.1\-8B\-Instruct is a text\-only LLM, we only ran language capability benchmarks against it\.
All in all, we aim to show three things through our experiments\. First, by evaluating our trained MLLMs on classification and captioning benchmarks, we quantitatively show that projector\-only\-trained MLLMs obtain 3D understanding scores comparable to existing methods\. Second, by comparing classification accuracy scores over time, we show that there is generally no evidence that joint training outperforms projector\-only training at any GPU\-hour threshold\. We validate this up to 16 hours, at which point we stop training\. Finally, we show through evaluating jointly trained backbones and base LLMs against standard benchmarks that joint training leads to drift in language, vision, and spatial reasoning capabilities to varying degrees\.
## 4Evaluations and Results
We report and analyze our 3D understanding and LM backbone benchmark results in this section\.
### 4\.13D Understanding
We report our 3D classification and captioning benchmark results for all Table[1](https://arxiv.org/html/2608.19726#S3.T1)variants to compare projector\-only training against existing models and against joint training\.
#### 4\.1\.1Competitive With Existing Methods
Figure[2](https://arxiv.org/html/2608.19726#S4.F2)compares projector\-only\-trained MLLMs against the baseline PointLLM models on 3D classification and captioning tasks\. We independently run these evaluations on existing models as our judge LLM differs from past papers\. We do not report PointLLM\-V2 scores due to the lack of publicly available weights\. Tables[2](https://arxiv.org/html/2608.19726#S4.T2)and[3](https://arxiv.org/html/2608.19726#S4.T3)show the exact scores for classification and captioning respectively, and additional tables in Appendix[C](https://arxiv.org/html/2608.19726#A3)include the scores of more existing models\.
Figure 2:Evaluation results of projector\-only\-trained MLLMs against baseline PointLLM models\. Chart \(a\) shows generative 3D object classification results on ModelNet40 \(M40\.\), Objaverse \(Obj\.\), and OmniObject3D \(Omni\.\) objects under a zero\-shot setting\. There are two prompt types per benchmark: an instruction\-style prompt \(I, “What is this?”\) and a completion\-style prompt \(C, “This is an object of”\)\. Each entry reports accuracy judged by GPT\-5\.6 Luna\[[28](https://arxiv.org/html/2608.19726#bib.bib29)\]\. Chart \(b\) shows the aggregate precision score on 3D object captioning tasks under the three judge LLMs\. More details are found in Appendix[B](https://arxiv.org/html/2608.19726#A2)\.All projector\-only\-trained MLLMs achieve scores comparable to or higher than those of PointLLM\[[43](https://arxiv.org/html/2608.19726#bib.bib2)\]\. In fact, all PointLLM scores lie below the 50th\{\}^\{\\text\{th\}\}percentile of each evaluation vertical\. In Appendix[C](https://arxiv.org/html/2608.19726#A3), projector\-only scores are also competitive when compared against improved methods, namely PointLLM\-R and MiniGPT\-3D\[[3](https://arxiv.org/html/2608.19726#bib.bib22),[37](https://arxiv.org/html/2608.19726#bib.bib24)\]\. The most fair comparison against the PointLLM family would be one made with P\-Llama8B, given that the architecture is the exact replica of that of PointLLM\-V2\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\]; our checkpoint outperforms both PointLLM models conclusively\.
Comparing between projector\-only backbones, we observe different trends across both task types\. While classification scores tend to improve with increasing backbone parameter count, higher captioning scores are obtained by MLLMs with native VLM backbones \(Qwen family\) rather than the text\-only LLM backbone \(Llama\)\. We hypothesize that this is because the VLM backbones have been trained on more data that resemble captions, allowing them to output generations more favorably scored by the judge LLMs for captioning tasks\.
Table 2:Generative 3D object classification results on ModelNet40 \(M40\.\), Objaverse \(Obj\.\), and OmniObject3D \(Omni\.\) under a zero\-shot setting as reported in Figure[2](https://arxiv.org/html/2608.19726#S4.F2)\(a\)\.Table 3:3D object captioning results on Objaverse as reported in Figure[2](https://arxiv.org/html/2608.19726#S4.F2)\(b\)\. C refers to correctness, H to hallucination, and P to precision \(aggregate of C and H\)\.Curriculum learning ablation\.In our training runs, all data is mixed together and sampled without discriminating between Stage 1 and Stage 2\. As an ablation, we trained an MLLM first on Stage 1 data and then on Stage 2 data, with an otherwise equivalent configuration to P\-Qwen4B\. This ablation variant is trained for the full 16 hours, with the proportion of Stage 1 data to Stage 2 data equal to that of all other variants\. The scores for this ablation are reported in extended Tables[9](https://arxiv.org/html/2608.19726#A3.T9)and[10](https://arxiv.org/html/2608.19726#A3.T10)\. Comparing these results to those of P\-Qwen4B, mixing data leads to better performance in some benchmarks but worse in others, thus no definitive evidence exists to argue for nor against it\.
Removing trained LoRA adapters from jointly trained MLLMs\.We also ran evaluations against jointly trained MLLMs without the trained LoRA adapters loaded\. Using only the trained projectors with the base LLM, the scores generally degrade by varying degrees, as reported in extended Tables[9](https://arxiv.org/html/2608.19726#A3.T9)and[10](https://arxiv.org/html/2608.19726#A3.T10)\. These degraded scores are lower than the corresponding projector\-only\-trained scores, demonstrating that the drift in existing language model capability reported in Section[4\.2](https://arxiv.org/html/2608.19726#S4.SS2)is a necessary tradeoff for jointly trained MLLMs to be competitive with projector\-only\-trained MLLMs in 3D understanding capability\. Interestingly, some scores still exceed PointLLM’s, which suggests that the projector carries much of the 3D understanding capability even during joint training\.
Through these results, we quantitatively demonstrate that projector\-only training leads to competitive 3D understanding capabilities\. This is also an important finding from the perspective of comparing projector\-only training against joint training \(i\.e\., comparing P\-\* against J\-\* variants\): we now establish that any findings regarding this comparison cannot be attributed to under\-training\.
#### 4\.1\.2Competitive With Joint Training
Figure[3](https://arxiv.org/html/2608.19726#S4.F3)plots accuracy scores of two 3D classification benchmarks against training GPU hours, specifically the averages of the I and C variants of the ModelNet40 and Objaverse benchmarks from Table[2](https://arxiv.org/html/2608.19726#S4.T2)\. We collect evaluation data for all variants at various GPU\-hour thresholds, namely 2, 4, 8, 12, and 16 hours\. As a companion, Figure[4](https://arxiv.org/html/2608.19726#S4.F4)plots accuracy scores against number of training samples seen, where each plotted point corresponds to the respective GPU\-hour threshold\. These figures compare projector\-only and joint training performance across different language model backbones, with exact scores reported in Table[11](https://arxiv.org/html/2608.19726#A3.T11)\.


Figure 3:M40\. \(top\) and Obj\. \(bottom\) mean accuracy against GPU hours\. Projector\-only training generally scores higher than joint training across all GPU hours\.

Figure 4:M40\. \(top\) and Obj\. \(bottom\) mean accuracy against training samples\. Plotted points in these charts correspond to the same GPU\-hour thresholds that in turn correspond to the plotted points in Figure[3](https://arxiv.org/html/2608.19726#S4.F3)\. Projector\-only training sees more samples than joint training at each GPU\-hour threshold as it is approximately twice as fast\.Generally, projector\-only training scores higher at most thresholds, with the difference in accuracy scores being negligible at several time thresholds\. Especially for M40\., projector\-only consistently outperforms joint\. These results establish that, under similar conditions, projector\-only training reaches competitive 3D understanding across all three backbones\. Joint training is therefore not necessary for 3D understanding to emerge\.
Since projector\-only training runs at approximately double the speed of joint training, it sees more samples within the same wall\-clock time\. Training speed also depends on the forward pass cost through the backbone, which accounts for the differences in sample counts across backbones but cancels within each P/J pair\. Reading Figure[4](https://arxiv.org/html/2608.19726#S4.F4)at matched sample counts, projector\-only training is more sample efficient when evaluated under M40\., but less sample efficient under Obj\. However, such as for P\-Llama8B on Obj\., projector\-only accuracy is still rising even as its joint counterpart stops improving, so part of joint training’s apparent sample advantage reflects earlier saturation rather than better learning\. In any case, sample efficiency is secondary in practice since a practitioner with a fixed GPU budget is constrained by compute rather than by samples, given that neither regime exhausted the training pool in our runs\.
We propose a hypothesis regarding the reason for the outlier trend in P\-Qwen4B and J\-Qwen4B on Obj\. In Figures[3](https://arxiv.org/html/2608.19726#S4.F3)and[4](https://arxiv.org/html/2608.19726#S4.F4), it is quite visible that both P\-Qwen4B and J\-Qwen4B have declined in performance prior to the 16\-hour mark, suggesting that there could be a better hyperparameter combination \(e\.g\., different seed for the dataloader\) that could have prevented early saturation and led to a narrower difference in performance between projector\-only and joint\. Having said that, it is interesting to see that J\-Qwen4B outperforms both J\-Llama8B and J\-Qwen9B while P\-Qwen4B lags behind both P\-Llama8B and P\-Qwen9B\.
### 4\.2Existing Language, Vision, and Spatial Reasoning Capabilities
Catastrophic forgetting in MLLMs is a known phenomenon\[[49](https://arxiv.org/html/2608.19726#bib.bib12)\]\. In this paper, none of our backbones have native 3D capabilities in their base versions as they are not equipped with 3D encoders; Qwen3\.5\-4B and Qwen3\.5\-9B are VLMs, while Llama\-3\.1\-8B\-Instruct is text\-only\. This means that catastrophic forgetting cannot occur for the 3D modality, rather it is only relevant for existing capabilities in language, vision, and image\-based spatial reasoning\. Since the projector maps solely features from the 3D point cloud encoder, and the LM backbone in projector\-only training is unmodified, catastrophic forgetting in this work only applies to joint training\.
To determine the extent of drift in existing capabilities, we evaluate the backbones of our jointly trained MLLMs at the 16\-hour checkpoint against standard language, vision, and spatial reasoning benchmarks\. We also evaluate the backbones of projector\-only\-trained MLLMs, which are equivalent to their base LLMs that have not been fine\-tuned on our dataset\. We report these results in Table[4](https://arxiv.org/html/2608.19726#S4.T4)\.
Table 4:Results for language, vision, and spatial reasoning benchmarks\. The LoRA fine\-tuned backbones are compared against the corresponding base backbone\.For language capability, there is an overall degradation in performance, especially for the Llama backbone\. Analyzing the model outputs, we find the significant drift for the Llama backbone is due to its inability to form coherent text for open\-ended questions, which is the format of IFEval\[[51](https://arxiv.org/html/2608.19726#bib.bib39)\], IFBench\[[29](https://arxiv.org/html/2608.19726#bib.bib40)\], and GSM8K\[[8](https://arxiv.org/html/2608.19726#bib.bib41)\]\. The Llama backbone also fails to follow instructions to generate code, hence collapsing to a zero score on HumanEval\[[7](https://arxiv.org/html/2608.19726#bib.bib42)\]\.
There is interestingly a general improvement in vision scores, relevant only for the Qwen backbones, though this quantitative result is misleading\. The benchmarks for which the scores noticeably improved, namely MMMU\[[47](https://arxiv.org/html/2608.19726#bib.bib46)\], MMMU\-Pro\[[48](https://arxiv.org/html/2608.19726#bib.bib37)\], MMMU\-Pro Vision\[[48](https://arxiv.org/html/2608.19726#bib.bib37)\], and MMStar\[[6](https://arxiv.org/html/2608.19726#bib.bib47)\], all contain multiple\-choice questions that are scored by renormalizing the model’s likelihoods over only the answer options\. This means the evaluations do not rely on decoded text generations, and that the resulting score is invariant to the absolute probability the model places on the correct answer\. Investigating the outputs to open\-ended benchmarks like BabyVision\[[5](https://arxiv.org/html/2608.19726#bib.bib34)\]and RealWorldQA\[[42](https://arxiv.org/html/2608.19726#bib.bib48)\], we see instruction\-following failure in jointly trained backbones \(e\.g\., responding to a question asking for a number with a caption\-like answer\) that may have arisen due to the narrow distribution of text types in the fine\-tuning 3D dataset\. These findings indicate that vision capability has not improved, and that the improved scores are actually misleading due to the mechanism of the evaluation\.
The same conclusion can be drawn from the spatial results\. Note that these benchmarks do not natively feed 3D inputs into the models; rather, they use 2D images paired with questions that are subcategorized as spatial reasoning questions\. Jointly trained backbones again often fail to follow instructions appropriately, such as responding with a description of the object rather than outputting the coordinates of the object of interest\. We again attribute this collapse to the dataset used to fine\-tune the MLLMs for 3D understanding\.
Overall, the degradation in the existing capabilities of the LM backbones is significant enough to dissuade practitioners from conducting joint training to build general MLLMs\. Since projector\-only training evidently achieves competitive 3D understanding capabilities, the effort to optimize hyperparameters for joint training to prevent backbone drifts may not be warranted\.
## 5Conclusion
In this work, we show that adapting the language model in training multimodal large language models is unnecessary\. By conducting training runs and evaluations across different language model backbones, we show that projector\-only training results in 3D MLLMs with capabilities comparable to those of existing baseline 3D MLLMs, and that projector\-only training is also competitive with joint training under matched compute budgets\. Hence, training the projector alone is sufficient for building an MLLM\.
Projector\-only training provides two demonstrable advantages over joint training\. Firstly, projector\-only training has a higher iteration speed than joint training, allowing it to see more training samples during the same wall time budget\. Secondly, as the LM backbone in projector\-only training is frozen, it experiences absolutely no drift in its existing capabilities\. Through relevant benchmarks, we show that joint training, on the other hand, does cause a noticeable degradation in existing capabilities of the LM backbone due to supervised fine\-tuning\.
As an attempt to understand why projector\-only training is so effective, we draw parallels to prompt engineering\. We hypothesize that optimizing the output of the projector is akin to searching for the best sequence of discrete tokens that yields the best performance in a large language model for a particular task\. Despite freezing the backbone, the MLLM learns to answer 3D understanding questions correctly, which means that the projector uses its input features from the encoder to elicit correct token distributions from the backbone\. We leave interpretability to future work\.
We are excited about the possibility of a more scalable, modular approach to training general MLLMs for all modalities with projector\-only training\. Since the LM backbone need not be fine\-tuned, the same LM can serve as the backbone for multiple learned projectors and their corresponding encoders for multiple modalities, each being able to be independently trained\. While our work is limited to 3D point cloud encoders, there has been prior work contrasting catastrophic forgetting between linear \(projector\-only\) and LoRA fine\-tuning that implicitly shows that projector\-only fine\-tuning still leads to 2D image understanding capability\[[49](https://arxiv.org/html/2608.19726#bib.bib12)\]\. We believe that testing projector\-only training on MLLMs of different modalities is a valuable future research direction towards modular MLLM training\.
## Acknowledgments and Disclosure of Funding
Nyx Iskandar and Saathvik Selvan completed this work as Research Engineer and Research Fellow at Ramen VR, respectively\. Slater Victoroff is the advisor of the project\. The authors would like to thank Andy Tsen and the Ramen VR team for their support and funding\.
## References
- \[1\]J\. Alayrac, J\. Donahue, P\. Luc, A\. Miech, I\. Barr, Y\. Hasson, K\. Lenc, A\. Mensch, K\. Millican, M\. Reynolds, R\. Ring, E\. Rutherford, S\. Cabi, T\. Han, Z\. Gong, S\. Samangooei, M\. Monteiro, J\. Menick, S\. Borgeaud, A\. Brock, A\. Nematzadeh, S\. Sharifzadeh, M\. Binkowski, R\. Barreira, O\. Vinyals, A\. Zisserman, and K\. Simonyan\(2022\)Flamingo: a visual language model for few\-shot learning\.External Links:2204\.14198,[Link](https://arxiv.org/abs/2204.14198)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[2\]Anthropic\(2025\)Introducing claude haiku 4\.5\.External Links:[Link](https://www.anthropic.com/news/claude-haiku-4-5)Cited by:[§B\.2](https://arxiv.org/html/2608.19726#A2.SS2.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1)\.
- \[3\]C\. Chen, Q\. Xu, W\. Zhou, and H\. Huang\(2026\)PointLLM\-r: enhancing 3d point cloud reasoning via chain\-of\-thought\.External Links:2605\.22013,[Link](https://arxiv.org/abs/2605.22013)Cited by:[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.7.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.6.1),[§1](https://arxiv.org/html/2608.19726#S1.p2.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1),[§4\.1\.1](https://arxiv.org/html/2608.19726#S4.SS1.SSS1.p2.1)\.
- \[4\]J\. Chen, Z\. Xu, X\. Pan, Y\. Hu, C\. Qin, T\. Goldstein, L\. Huang, T\. Zhou, S\. Xie, S\. Savarese, L\. Xue, C\. Xiong, and R\. Xu\(2025\)BLIP3\-o: a family of fully open unified multimodal models\-architecture, training and dataset\.External Links:2505\.09568,[Link](https://arxiv.org/abs/2505.09568)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[5\]L\. Chen, W\. Xie, Y\. Liang, H\. He, H\. Zhao, Z\. Yang, Z\. Huang, H\. Wu, H\. Lu, Y\. charles, Y\. Bao, Y\. Fan, G\. Li, H\. Shen, X\. Chen, W\. Xu, S\. Si, Z\. Cai, W\. Chai, Z\. Huang, F\. Liu, T\. Liu, B\. Chang, M\. Wu, X\. Hu, K\. Chen, Y\. Ren, Y\. Liu, Y\. Gong, and K\. Li\(2026\)BabyVision: visual reasoning beyond language\.External Links:2601\.06521,[Link](https://arxiv.org/abs/2601.06521)Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.18.1)\.
- \[6\]L\. Chen, J\. Li, X\. Dong, P\. Zhang, Y\. Zang, Z\. Chen, H\. Duan, J\. Wang, Y\. Qiao, D\. Lin,et al\.\(2024\)Are we on the right way for evaluating large vision\-language models?\.arXiv preprint arXiv:2403\.20330\.Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.17.1)\.
- \[7\]M\. Chen, J\. Tworek, H\. Jun, Q\. Yuan, H\. P\. de Oliveira Pinto, J\. Kaplan, H\. Edwards, Y\. Burda, N\. Joseph, G\. Brockman, A\. Ray, R\. Puri, G\. Krueger, M\. Petrov, H\. Khlaaf, G\. Sastry, P\. Mishkin, B\. Chan, S\. Gray, N\. Ryder, M\. Pavlov, A\. Power, L\. Kaiser, M\. Bavarian, C\. Winter, P\. Tillet, F\. P\. Such, D\. Cummings, M\. Plappert, F\. Chantzis, E\. Barnes, A\. Herbert\-Voss, W\. H\. Guss, A\. Nichol, A\. Paino, N\. Tezak, J\. Tang, I\. Babuschkin, S\. Balaji, S\. Jain, W\. Saunders, C\. Hesse, A\. N\. Carr, J\. Leike, J\. Achiam, V\. Misra, E\. Morikawa, A\. Radford, M\. Knight, M\. Brundage, M\. Murati, K\. Mayer, P\. Welinder, B\. McGrew, D\. Amodei, S\. McCandlish, I\. Sutskever, and W\. Zaremba\(2021\)Evaluating large language models trained on code\.External Links:2107\.03374Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p3.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.12.1)\.
- \[8\]K\. Cobbe, V\. Kosaraju, M\. Bavarian, M\. Chen, H\. Jun, L\. Kaiser, M\. Plappert, J\. Tworek, J\. Hilton, R\. Nakano, C\. Hesse, and J\. Schulman\(2021\)Training verifiers to solve math word problems\.arXiv preprint arXiv:2110\.14168\.Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p3.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.9.1)\.
- \[9\]W\. Dai, J\. Li, D\. Li, A\. M\. H\. Tiong, J\. Zhao, W\. Wang, B\. Li, P\. Fung, and S\. Hoi\(2023\)InstructBLIP: towards general\-purpose vision\-language models with instruction tuning\.External Links:2305\.06500,[Link](https://arxiv.org/abs/2305.06500)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[10\]M\. Deitke, R\. Liu, M\. Wallingford, H\. Ngo, O\. Michel, A\. Kusupati, A\. Fan, C\. Laforte, V\. Voleti, S\. Y\. Gadre, E\. VanderBilt, A\. Kembhavi, C\. Vondrick, G\. Gkioxari, K\. Ehsani, L\. Schmidt, and A\. Farhadi\(2023\)Objaverse\-xl: a universe of 10m\+ 3d objects\.External Links:2307\.05663,[Link](https://arxiv.org/abs/2307.05663)Cited by:[§A\.2](https://arxiv.org/html/2608.19726#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1)\.
- \[11\]M\. Deitke, D\. Schwenk, J\. Salvador, L\. Weihs, O\. Michel, E\. VanderBilt, L\. Schmidt, K\. Ehsani, A\. Kembhavi, and A\. Farhadi\(2022\)Objaverse: a universe of annotated 3d objects\.arXiv preprint arXiv:2212\.08051\.Cited by:[§A\.2](https://arxiv.org/html/2608.19726#A1.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p3.1),[§3\.2](https://arxiv.org/html/2608.19726#S3.SS2.p2.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1)\.
- \[12\]J\. Devlin, M\. Chang, K\. Lee, and K\. Toutanova\(2019\)BERT: pre\-training of deep bidirectional transformers for language understanding\.External Links:1810\.04805,[Link](https://arxiv.org/abs/1810.04805)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p2.1)\.
- \[13\]D\. Driess, F\. Xia, M\. S\. M\. Sajjadi, C\. Lynch, A\. Chowdhery, B\. Ichter, A\. Wahid, J\. Tompson, Q\. Vuong, T\. Yu, W\. Huang, Y\. Chebotar, P\. Sermanet, D\. Duckworth, S\. Levine, V\. Vanhoucke, K\. Hausman, M\. Toussaint, K\. Greff, A\. Zeng, I\. Mordatch, and P\. Florence\(2023\)PaLM\-e: an embodied multimodal language model\.External Links:2303\.03378,[Link](https://arxiv.org/abs/2303.03378)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[14\]M\. Du, B\. Wu, Z\. Li, X\. Huang, and Z\. Wei\(2024\)EmbSpatial\-bench: benchmarking spatial understanding for embodied tasks with large vision\-language models\.External Links:2406\.05756,[Link](https://arxiv.org/abs/2406.05756)Cited by:[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.22.1)\.
- \[15\]A\. P\. Gema, J\. O\. J\. Leang, G\. Hong, A\. Devoto, A\. C\. M\. Mancino, R\. Saxena, X\. He, Y\. Zhao, X\. Du, M\. R\. G\. Madani, C\. Barale, R\. McHardy, J\. Harris, J\. Kaddour, E\. van Krieken, and P\. Minervini\(2025\)Are we done with mmlu?\.External Links:2406\.04127,[Link](https://arxiv.org/abs/2406.04127)Cited by:[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.5.1)\.
- \[16\]Gemini Robotics Team\(2025\)Gemini robotics: bringing ai into the physical world\.External Links:2503\.20020,[Link](https://arxiv.org/abs/2503.20020)Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.21.1)\.
- \[17\]Google\(2026\)Introducing gemini 3\.6 flash, 3\.5 flash\-lite, and 3\.5 flash cyber\.Google\.External Links:[Link](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/)Cited by:[§B\.2](https://arxiv.org/html/2608.19726#A2.SS2.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1)\.
- \[18\]D\. Hendrycks and K\. Gimpel\(2023\)Gaussian error linear units \(gelus\)\.External Links:1606\.08415,[Link](https://arxiv.org/abs/1606.08415)Cited by:[§3\.1](https://arxiv.org/html/2608.19726#S3.SS1.p3.1)\.
- \[19\]E\. J\. Hu, Y\. Shen, P\. Wallis, Z\. Allen\-Zhu, Y\. Li, S\. Wang, L\. Wang, and W\. Chen\(2021\)LoRA: low\-rank adaptation of large language models\.External Links:2106\.09685,[Link](https://arxiv.org/abs/2106.09685)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p2.1),[§3\.2](https://arxiv.org/html/2608.19726#S3.SS2.p2.1)\.
- \[20\]R\. Huang, M\. Li, D\. Yang, J\. Shi, X\. Chang, Z\. Ye, Y\. Wu, Z\. Hong, J\. Huang, J\. Liu, Y\. Ren, Z\. Zhao, and S\. Watanabe\(2023\)AudioGPT: understanding and generating speech, music, sound, and talking head\.External Links:2304\.12995,[Link](https://arxiv.org/abs/2304.12995)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[21\]S\. Huang, L\. Dong, W\. Wang, Y\. Hao, S\. Singhal, S\. Ma, T\. Lv, L\. Cui, O\. K\. Mohammed, B\. Patra, Q\. Liu, K\. Aggarwal, Z\. Chi, J\. Bjorck, V\. Chaudhary, S\. Som, X\. Song, and F\. Wei\(2023\)Language is not all you need: aligning perception with language models\.External Links:2302\.14045,[Link](https://arxiv.org/abs/2302.14045)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[22\]H\. Liu, C\. Li, Q\. Wu, and Y\. J\. Lee\(2023\)Visual instruction tuning\.InAdvances in Neural Information Processing Systems,A\. Oh, T\. Naumann, A\. Globerson, K\. Saenko, M\. Hardt, and S\. Levine \(Eds\.\),Vol\.36,pp\. 34892–34916\.External Links:[Document](https://dx.doi.org/10.52202/075280-1516),[Link](https://proceedings.neurips.cc/paper_files/paper/2023/file/6dcf277ea32ce3288914faf369fe6de0-Paper-Conference.pdf)Cited by:[§1](https://arxiv.org/html/2608.19726#S1.p1.1)\.
- \[23\]Llama Team\(2024\)The llama 3 herd of models\.External Links:2407\.21783,[Link](https://arxiv.org/abs/2407.21783)Cited by:[§3\.1](https://arxiv.org/html/2608.19726#S3.SS1.p4.1)\.
- \[24\]T\. Luo, C\. Rockwell, H\. Lee, and J\. Johnson\(2023\)Scalable 3d captioning with pretrained models\.External Links:2306\.07279,[Link](https://arxiv.org/abs/2306.07279)Cited by:[§B\.2](https://arxiv.org/html/2608.19726#A2.SS2.SSS0.Px1.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p3.1)\.
- \[25\]A\. Marcu, L\. Chen, J\. Hünermann, A\. Karnsund, B\. Hanotte, P\. Chidananda, S\. Nair, V\. Badrinarayanan, A\. Kendall, J\. Shotton, and O\. Sinavski\(2023\)LingoQA: visual question answering for autonomous driving\.arXiv preprint arXiv:2312\.14115\.Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.24.1)\.
- \[26\]T\. Mihaylov, P\. Clark, T\. Khot, and A\. Sabharwal\(2018\)Can a suit of armor conduct electricity? a new dataset for open book question answering\.InEMNLP,Cited by:[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.11.1)\.
- \[27\]OpenAI\(2024\)Hello gpt\-4o \| openai\.External Links:[Link](https://openai.com/index/hello-gpt-4o/)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p3.1)\.
- \[28\]OpenAI\(2026\)GPT\-5\.6: frontier intelligence that scales with your ambition \| openai\.External Links:[Link](https://openai.com/index/gpt-5-6/)Cited by:[§B\.1](https://arxiv.org/html/2608.19726#A2.SS1.SSS0.Px5.p1.1),[§B\.2](https://arxiv.org/html/2608.19726#A2.SS2.SSS0.Px4.p1.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1),[Figure 2](https://arxiv.org/html/2608.19726#S4.F2)\.
- \[29\]V\. Pyatkin, S\. Malik, V\. Graf, H\. Ivison, S\. Huang, P\. Dasigi, N\. Lambert, and H\. Hajishirzi\(20252025\)Generalizing verifiable instruction following\.Vol\.38\.Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p3.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.8.1)\.
- \[30\]C\. R\. Qi, H\. Su, K\. Mo, and L\. J\. Guibas\(2017\)PointNet: deep learning on point sets for 3d classification and segmentation\.External Links:1612\.00593,[Link](https://arxiv.org/abs/1612.00593)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p2.1)\.
- \[31\]Z\. Qi, R\. Dong, G\. Fan, Z\. Ge, X\. Zhang, K\. Ma, and L\. Yi\(2023\)Contrast with reconstruct: contrastive 3d representation learning guided by generative pretraining\.External Links:2302\.02318,[Link](https://arxiv.org/abs/2302.02318)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p1.1)\.
- \[32\]Z\. Qi, R\. Dong, S\. Zhang, H\. Geng, C\. Han, Z\. Ge, L\. Yi, and K\. Ma\(2024\)ShapeLLM: universal 3d object understanding for embodied interaction\.External Links:2402\.17766,[Link](https://arxiv.org/abs/2402.17766)Cited by:[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.3.1),[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.4.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.2.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.3.1),[§1](https://arxiv.org/html/2608.19726#S1.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1)\.
- \[33\]Qwen Team\(2026\)Qwen3\.5: towards native multimodal agents\.External Links:[Link](https://qwen.ai/blog?id=qwen3.5)Cited by:[§3\.1](https://arxiv.org/html/2608.19726#S3.SS1.p4.1)\.
- \[34\]A\. Radford, J\. Wu, R\. Child, D\. Luan, D\. Amodei, and I\. Sutskever\(2019\)Language models are unsupervised multitask learners\.OpenAI\.Note:Accessed: 2024\-11\-15External Links:[Link](https://cdn.openai.com/better-language-models/language_models_are_unsupervised_multitask_learners.pdf)Cited by:[§3\.2](https://arxiv.org/html/2608.19726#S3.SS2.p1.1)\.
- \[35\]D\. Rein, B\. L\. Hou, A\. C\. Stickland, J\. Petty, R\. Y\. Pang, J\. Dirani, J\. Michael, and S\. R\. Bowman\(2023\)GPQA: a graduate\-level google\-proof q&a benchmark\.External Links:2311\.12022,[Link](https://arxiv.org/abs/2311.12022)Cited by:[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.6.1)\.
- \[36\]K\. Sakaguchi, R\. L\. Bras, C\. Bhagavatula, and Y\. Choi\(2019\)WinoGrande: an adversarial winograd schema challenge at scale\.External Links:1907\.10641,[Link](https://arxiv.org/abs/1907.10641)Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.10.1)\.
- \[37\]Y\. Tang, X\. Han, X\. Li, Q\. Yu, Y\. Hao, L\. Hu, and M\. Chen\(2024\)MiniGPT\-3d: efficiently aligning 3d point clouds with large language models using 2d priors\.External Links:2405\.01413,[Link](https://arxiv.org/abs/2405.01413)Cited by:[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.8.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.7.1),[§1](https://arxiv.org/html/2608.19726#S1.p2.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§4\.1\.1](https://arxiv.org/html/2608.19726#S4.SS1.SSS1.p2.1)\.
- \[38\]A\. Vaswani, N\. Shazeer, N\. Parmar, J\. Uszkoreit, L\. Jones, A\. N\. Gomez, L\. Kaiser, and I\. Polosukhin\(2023\)Attention is all you need\.External Links:1706\.03762,[Link](https://arxiv.org/abs/1706.03762)Cited by:[§3\.1](https://arxiv.org/html/2608.19726#S3.SS1.p4.1)\.
- \[39\]Y\. Wang, X\. Ma, G\. Zhang, Y\. Ni, A\. Chandra, S\. Guo, W\. Ren, A\. Arulraj, X\. He, Z\. Jiang,et al\.\(2024\)Mmlu\-pro: a more robust and challenging multi\-task language understanding benchmark\.arXiv preprint arXiv:2406\.01574\.Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.4.1)\.
- \[40\]T\. Wu, J\. Zhang, X\. Fu, Y\. Wang, L\. P\. Jiawei Ren, W\. Wu, L\. Yang, J\. Wang, C\. Qian, D\. Lin, and Z\. Liu\(2023\)OmniObject3D: large\-vocabulary 3d object dataset for realistic perception, reconstruction and generation\.InIEEE/CVF Conference on Computer Vision and Pattern Recognition \(CVPR\),Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1)\.
- \[41\]Z\. Wu, S\. Song, A\. Khosla, F\. Yu, L\. Zhang, X\. Tang, and J\. Xiao\(2015\)3D shapenets: a deep representation for volumetric shapes\.External Links:1406\.5670,[Link](https://arxiv.org/abs/1406.5670)Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1)\.
- \[42\]xAI\(2024\)RealWorldQA: a benchmark for real\-world spatial understanding in multimodal models\.Hugging Face\.Note:[https://huggingface\.co/datasets/xai\-org/RealworldQA](https://huggingface.co/datasets/xai-org/RealworldQA)Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.19.1)\.
- \[43\]R\. Xu, X\. Wang, T\. Wang, Y\. Chen, J\. Pang, and D\. Lin\(2024\)PointLLM: empowering large language models to understand point clouds\.InECCV,Cited by:[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.5.1),[Table 10](https://arxiv.org/html/2608.19726#A3.T10.2.6.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.4.1),[Table 9](https://arxiv.org/html/2608.19726#A3.T9.2.5.1),[§1](https://arxiv.org/html/2608.19726#S1.p1.1),[§1](https://arxiv.org/html/2608.19726#S1.p2.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p3.1),[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p3.1),[§4\.1\.1](https://arxiv.org/html/2608.19726#S4.SS1.SSS1.p2.1),[Table 2](https://arxiv.org/html/2608.19726#S4.T2.2.2.1),[Table 2](https://arxiv.org/html/2608.19726#S4.T2.2.3.1),[Table 3](https://arxiv.org/html/2608.19726#S4.T3.2.3.1),[Table 3](https://arxiv.org/html/2608.19726#S4.T3.2.4.1)\.
- \[44\]R\. Xu, S\. Yang, X\. Wang, T\. Wang, Y\. Chen, J\. Pang, and D\. Lin\(2025\)PointLLM\-v2: empowering large language models to better understand point clouds\.IEEE Transactions on Pattern Analysis and Machine Intelligence\(\),pp\. 1–15\.External Links:[Document](https://dx.doi.org/10.1109/TPAMI.2025.3590784)Cited by:[§A\.2](https://arxiv.org/html/2608.19726#A1.SS2.SSS0.Px1.p1.1),[§A\.3](https://arxiv.org/html/2608.19726#A1.SS3.p1.1),[§1](https://arxiv.org/html/2608.19726#S1.p2.1),[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p3.1),[§3\.1](https://arxiv.org/html/2608.19726#S3.SS1.p3.1),[§3\.2](https://arxiv.org/html/2608.19726#S3.SS2.p2.1),[§4\.1\.1](https://arxiv.org/html/2608.19726#S4.SS1.SSS1.p2.1)\.
- \[45\]L\. Xue, N\. Yu, S\. Zhang, A\. Panagopoulou, J\. Li, R\. Martín\-Martín, J\. Wu, C\. Xiong, R\. Xu, J\. C\. Niebles, and S\. Savarese\(2024\)ULIP\-2: towards scalable multimodal pre\-training for 3d understanding\.External Links:2305\.08275,[Link](https://arxiv.org/abs/2305.08275)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p2.1)\.
- \[46\]X\. Yu, L\. Tang, Y\. Rao, T\. Huang, J\. Zhou, and J\. Lu\(2022\)Point\-bert: pre\-training 3d point cloud transformers with masked point modeling\.External Links:2111\.14819,[Link](https://arxiv.org/abs/2111.14819)Cited by:[§2](https://arxiv.org/html/2608.19726#S2.p1.1),[§2](https://arxiv.org/html/2608.19726#S2.p2.1)\.
- \[47\]X\. Yue, Y\. Ni, K\. Zhang, T\. Zheng, R\. Liu, G\. Zhang, S\. Stevens, D\. Jiang, W\. Ren, Y\. Sun, C\. Wei, B\. Yu, R\. Yuan, R\. Sun, M\. Yin, B\. Zheng, Z\. Yang, Y\. Liu, W\. Huang, H\. Sun, Y\. Su, and W\. Chen\(2024\)MMMU: a massive multi\-discipline multimodal understanding and reasoning benchmark for expert agi\.InProceedings of CVPR,Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.14.1)\.
- \[48\]X\. Yue, T\. Zheng, Y\. Ni, Y\. Wang, K\. Zhang, S\. Tong, Y\. Sun, B\. Yu, G\. Zhang, H\. Sun, Y\. Su, W\. Chen, and G\. Neubig\(2025\)MMMU\-pro: a more robust multi\-discipline multimodal understanding benchmark\.External Links:2409\.02813,[Link](https://arxiv.org/abs/2409.02813)Cited by:[§3\.3](https://arxiv.org/html/2608.19726#S3.SS3.p4.1),[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p4.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.15.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.16.1)\.
- \[49\]Y\. Zhai, S\. Tong, X\. Li, M\. Cai, Q\. Qu, Y\. J\. Lee, and Y\. Ma\(2023\)Investigating the catastrophic forgetting in multimodal large language models\.External Links:2309\.10313,[Link](https://arxiv.org/abs/2309.10313)Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p1.1),[§5](https://arxiv.org/html/2608.19726#S5.p4.1)\.
- \[50\]E\. Zhou, J\. An, C\. Chi, Y\. Han, S\. Rong, C\. Zhang, P\. Wang, Z\. Wang, T\. Huang, L\. Sheng, and S\. Zhang\(2026\)RoboRefer: towards spatial referring with reasoning in vision\-language models for robotics\.External Links:2506\.04308,[Link](https://arxiv.org/abs/2506.04308)Cited by:[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.23.1)\.
- \[51\]J\. Zhou, T\. Lu, S\. Mishra, S\. Brahma, S\. Basu, Y\. Luan, D\. Zhou, and L\. Hou\(2023\)Instruction\-following evaluation for large language models\.External Links:2311\.07911,[Link](https://arxiv.org/abs/2311.07911)Cited by:[§4\.2](https://arxiv.org/html/2608.19726#S4.SS2.p3.1),[Table 4](https://arxiv.org/html/2608.19726#S4.T4.2.7.1)\.
## Appendix ATraining Configuration
### A\.1Input Representation and Model
Each object enters the encoder asn=8192n=8192points withd=6d=6channels\. Positions are normalized to the unit sphere and RGB values are scaled to\[−1,1\]\[\-1,1\]\.
The projector is a 3\-layer MLP mappingc→1024→2048→c′c\\rightarrow 1024\\rightarrow 2048\\rightarrow c^\{\\prime\}with GELU applied between consecutive linear layers and no activations on the output\. The projector is trained and stored infloat32while the backbone runs inbfloat16, since the projector is small enough that full precision costs very little\. The two boundary tokens<\|pc3d\_start\|\>and<\|pc3d\_end\|\>are appended to the vocabulary and initialized to the mean of the base embedding matrix and frozen\. Since our models are not required to output these tokens, we are able to freeze the entire backbone during projector\-only training, allowing us to guarantee zero drift of the base LM capabilities\.
Table 5:Projector and LoRA\-adapter parameter counts, and projector input/output shapes, for each backbone\.
### A\.2Data and Sampling
##### Corpus construction\.
Training rows are drawn from the union of the Stage 1 and Stage 2 subsets of the PointLLM\-V2 dataset\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\], restricted to objects appearing in Objaverse\[[11](https://arxiv.org/html/2608.19726#bib.bib15)\]and excluding Objaverse\-XL\[[10](https://arxiv.org/html/2608.19726#bib.bib21)\]\. After filtering, the pool contains 1,368,302 rows over 661,375 unique objects, containing 661,375 Stage 1 rows and 706,927 Stage 2 rows\. Stage 1 contains exactly one brief description row per object, while Stage 2 contains several rows per object, divided into 77,440 captioning, 310,415 question\-answering, and 319,072 referring rows\.
##### Sampling\.
Rows are sampled uniformly without replacement using a fixed seed of 0 and consumed in a single fixed order shared across all variants, so that a projector\-only run and its joint counterpart see the exact same stream of data\.
##### Sequence construction\.
Prompts are formatted with one of the system messages shown in Table[6](https://arxiv.org/html/2608.19726#A1.T6), followed by point tokens inserted between<\|pc3d\_start\|\>and<\|pc3d\_end\|\>tags, and ending with a specific question from our training dataset\. The labels on the prompt tokens are masked with−\-100 so only the model response tokens contribute to our loss\. Sequences are truncated at 512 tokens, although this limit is never reached in practice\.
Table 6:The set of system message paraphrases, drawn uniformly at random during training\.
### A\.3Optimization
All parameters are optimized with AdamW under the token\-averaged loss objectiveℒ\(θ\)=−1T∑logPθ\(xt\+1\|x≤t\)\\mathcal\{L\}\(\\theta\)=\-\\frac\{1\}\{T\}\\sum\{\\log\{P\_\{\\theta\}\(x\_\{t\+1\}\|x\_\{\\leq t\}\)\}\}\. Projector\-only training optimizes a single parameter group while joint training optimizes two groups under one optimizer instance, with separate projector and LoRA learning rates\. The schedule is constant with no warmup and zero weight decay in both regimes\. As the paper primarily concerns comparisons between projector\-only and joint training regimes, we deliberately choose a plain config to avoid excessive hyperparameter tuning for the sake of maximizing performance\. The projector learning rate is the one setting that varies across configurations, chosen empirically using a small sweep\. Table[7](https://arxiv.org/html/2608.19726#A1.T7)shows all hyperparameter configurations; we generally followed the learning rates in PointLLM\-V2\[[44](https://arxiv.org/html/2608.19726#bib.bib1)\], though we modified the Llama8B learning rates as the original values led to underperforming MLLMs for both projector\-only and joint training\.
Table 7:Training and evaluation hyperparameters, shared across every configuration except the learning rates\.
### A\.4Throughput and Samples Seen
Table[8](https://arxiv.org/html/2608.19726#A1.T8)shows the total number of training samples seen by each MLLM variant after 16 GPU hours\. Given the same language model backbone, projector\-only training processes samples at approximately twice the rate of joint training\. This ratio holds near2×2\\timesacross all three pairings, and is due to joint training having to track and backpropagate into LoRA adapter weights\.
A secondary ordering is visible within each training regime, where each backbone has a different relative training rate\. Qwen3\.5\-4B is the fastest, followed by Llama\-3\.1\-8B\-Instruct and Qwen3\.5\-9B, as expected when the frozen forward pass dominates cost\.
Table 8:Total number of training samples seen by each MLLM variant, average throughput, and average step rate after 16 GPU hours\.
## Appendix BEvaluation Procedure
### B\.1Generative Zero\-Shot Classification
##### ModelNet40\.
We use PointLLM’s released test file, which contains the standard 2,468 object test split, already farthest\-point sampled to 8192 points\. We keep the XYZ channels and discard the surface normals\. ModelNet40 is geometry\-only CAD data with no color, so every point is assignedRGB=−1\\mathrm\{RGB\}=\-1\(black\) under our color convention\.
##### Objaverse\.
We use PointLLM’s 3,000 object held\-out benchmark with their brief\-description ground truth as the reference answer\. There is no fixed label set, since the reference is a short free\-form caption for each object\. The point clouds are the release’s own 8192\-point colored point clouds, extracted per object from the released data\.
##### OmniObject3D\.
The ground truth is PointLLM\-R’s released brief\-description validation set, which contains 5,900 objects over 216 categories\. The official OmniObject3D point clouds carry no color, so we transfer color onto them from the raw textured scans, keeping the point positions byte\-identical and changing only the color channels\.
##### Common processing\.
At evaluation time, each point cloud is subsampled to 8192 points, its coordinates are normalized to the unit sphere, and color is passed in\[−1,1\]\[\-1,1\]\. Generations are greedy with a 256\-token budget, and each benchmark is run under both variants of PointLLM’s classification prompts:"This is an object of"\(C\) and"What is this?"\(I\)\.
##### Judge protocol\.
Generations are scored by GPT\-5\.6 Luna\[[28](https://arxiv.org/html/2608.19726#bib.bib29)\]\(gpt\-5\.6\-luna\), with a 96\-token visible reply budget\. Two different rubrics are used, shown in Figures[5](https://arxiv.org/html/2608.19726#A2.F5)and[6](https://arxiv.org/html/2608.19726#A2.F6), depending on whether a fixed set of classes is available\. On Objaverse and OmniObject3D, the judge sees the reference caption and simply decides whether there is a match\. On ModelNet40, however, the judge maps the free\-form generation onto one of the 40 classes and returns the closest one, without seeing the ground truth answer\.
### B\.2Object Captioning
##### Evaluation set\.
We use the same 3,000 Objaverse objects as the classification benchmark, each paired with the 8 rendered views shipped by Cap3D\[[24](https://arxiv.org/html/2608.19726#bib.bib9)\]\.
##### Model prompt\.
We use PointLLM’s captioning prompt verbatim: “Caption this 3D model in detail\.” Captions are decoded greedily with a 256\-token budget, identical to the classification evaluations\.
##### Scoring rubric\.
Each caption is scored against the 8 views by an LLM judge with no reference caption; the images are the only ground truth\. The judge identifies the key visual aspects of the object \(category, color, shape, usage, material\) and returns a correctness scoreCCand a hallucination scoreHHas*point counts*, not percentages: one correctness point per correctly stated aspect \(fractional credit in\[0,1\]\[0,1\]for partial matches\), and one hallucination point per aspect the caption asserts that the images contradict\. Precision is computed similarly to PointLLM\-V2, as a ratio of the total correctness score to the sum of the correctness and hallucination scores over the entire evaluation set\.
##### Judges\.
We use three frontier LLMs as judges: GPT\-5\.6 Luna\[[28](https://arxiv.org/html/2608.19726#bib.bib29)\]\(gpt\-5\.6\-luna\), Claude Haiku 4\.5\[[2](https://arxiv.org/html/2608.19726#bib.bib30)\]\(claude\-haiku\-4\-5\-20251001\), and Gemini 3\.5 Flash\-Lite\[[17](https://arxiv.org/html/2608.19726#bib.bib31)\]\(gemini\-3\.5\-flash\-lite\), each with a 200\-token visible reply budget\. To account for variance between these judges, we average their precision scores together, as is done in PointLLM\-V2\. All three judges receive an identical single user message \(no system message\), shown in Figure[7](https://arxiv.org/html/2608.19726#A2.F7)\.
Figure 5:The judge prompt for generative zero\-shot classification on Objaverse and OmniObject3D\. Placeholders in braces are substituted at evaluation time; \{gold\} and \{generation\} refer to the ground\-truth caption and the model’s response, respectively\.Analyze two sentences and determine if they’re referring to the same general object or concept, focusing on the type of object, not attributes such as color, size, or shape\. Respond with ’T’ if they refer to the same thing and ’F’ if not\. Also, provide a brief rationale \(no more than 20 words\) for your judgment\.
Example:
Input: 1\. A black and brown colored gun\. 2\. The 3D object is a representation of a futuristic, high\-tech gun crafted from a glossy black material\. Distinctive features include its metallic handrail, giving an impression of a robust mechanized design\. The gun, possibly used in a sci\-fi or futuristic setting, denotes advanced technology and might include functionalities such as voice recognition, aiming systems, or biometric triggers\.
Output: T\#Both refer to a gun\.
Input: 1\. A yellow and white fish with black stripes and fins\. 2\. This is a 3D model of a vibrant, polka\-dotted toy fish that is predominantly orange on the body, shifting to white on the belly\. The toy has dark brown spots that enhance its appearance, potentially mimicking the natural patterns found on real\-life fish\. It’s an ideal object for educational purposes, helping to introduce children to marine life, as well as serving as a playful item in a playroom or nursery\.
Output: T\#Both refer to a fish\.
Input: 1\. A white cartoon scorpion with eight legs\. 2\. This is a 3D object model representing a cartoon version of a rare type of spider\. The entire model is rendered in white, which highlights its unique and exaggerated characteristics such as multiple legs and a funnel\-like body\. Its cartoonish appeal makes it more appealing to a younger audience, and it could possibly be used in animations or educational materials to teach children about spiders in a less intimidating way\.
Output: F\#One is a scorpion and the other is a spider\.
Now, analyze the following:
Input: 1\. \{gold\} 2\. \{generation\}
Output:Figure 6:The judge prompt for generative zero\-shot classification on ModelNet40\. The placeholder \{generation\} refers to the model’s response\.Given the following free\-form description of a 3D object, please determine the most probable class index from the following 40 available categories, even if the description doesn’t clearly refer to any one of them\. Make your best\-educated guess based on the information provided\. If the description already contains a valid index, then the index should be selected\. If it contains more than one valid index, then randomly select one index \(specify your reason\)\. If there is no valid index and it cannot be inferred from the information, return “\-1\#NA\#Cannot infer”\.
Categories:
0: airplane 1: bathtub 2: bed 3: bench 4: bookshelf 5: bottle 6: bowl 7: car 8: chair 9: cone 10: cup 11: curtain 12: desk 13: door 14: dresser 15: flower\_pot 16: glass\_box 17: guitar 18: keyboard 19: lamp 20: laptop 21: mantel 22: monitor 23: night\_stand 24: person 25: piano 26: plant 27: radio 28: range\_hood 29: sink 30: sofa 31: stairs 32: stool 33: table 34: tent 35: toilet 36: tv\_stand 37: vase 38: wardrobe 39: xbox
Reply with the format of “index\#class\#short reason \(no more than 10 words\)”\.
Examples:
Input: This is a 3D object model of a cartoon white truck\.
Output: 7\#car\#Closest match to “car” in categories\.
Input: A green leaf in a flower pot\.
Output: 26\#plant\#The primary subject “leaf” directly indicates a plant\.
Input: It’s difficult to determine the exact type of this object due to insufficient details\. But it seems to be like a piece of furniture\.
Output: 33\#table\#Randomly select one kind of furniture from the list\.
Input: I cannot determine the specific type of the object without additional information or context\.
Output: \-1\#NA\#Cannot infer\.
Now analyze the following:
Input: \{generation\}
Output:Figure 7:The judge prompt for object captioning on Objaverse\. Placeholders in braces are substituted at evaluation time; \{generation\} refers to the model’s generated caption\. The judge additionally receives 8 rendered views of the object alongside this text\.You are provided with 8 images of an object taken from different angles\. Your task is to evaluate a model\-generated caption based on the visual information in these images\. Identify the key aspects \(e\.g\., category, color, shape, usage, material\) from the images and calculate the percentage of these aspects that are correctly mentioned or partially matched in the model\-generated caption\. Assign correctness points for each distinct correct attribute\. Partial correctness should be graded on a scale of 0 to 1 depending on accuracy, with similar concepts considered for partial scores\. Each aspect contributes equally to the final score, with no extra penalties for repeated inaccuracies of the same attribute\. Also, assign hallucination points for incorrect details in the model output, with one point per incorrect attribute\. Repetitive inaccuracies based on one attribute should incur only a single hallucination point\. However, if a detail is described uncertainly or speculatively \(e\.g\., using phrases like ’probably for’, ’possibly’, ’resemble’ or ’indicating’\), assign fewer or no hallucination points\. Provide your score and a short justification \(less than 25 words\) in the format of “Output: score\#your reason”
Examples \(corresponding images omitted\):
Example 1
Model: The object is a geometric shape with a complex design featuring interlocking blue and white elements\. It appears to be a three\-dimensional structure with angular and rectangular cutouts, creating a visually intricate form\. The blue elements form the primary color scheme, contrasted by white areas that highlight the depth and complexity of the design\.
Correctness: 4\.5\#Correct identification of geometric shape, colors, angular and rectangular cutouts, and three\-dimensional structure\.
Hallucination: 0\#No incorrect details\.
Example 2
Model: The object is a green, cartoonish character with a rounded body and two short limbs protruding from the sides\. It has two antennae on top of its head and a pair of eyes on the front\. The overall shape is smooth and symmetrical, resembling a playful or animated figure\.
Correctness: 3\.5\#The object is green, cartoonish, and has a rounded body with two limbs\.
Hallucination: 2\.0\#No antennae or eyes are visible, shape is not entirely smooth\.
Model: \{generation\}
Now give your score and reason without adding extra content\. Reply with exactly two lines, and express both scores as point counts \(a number of attributes, decimals allowed for partial credit\) — not percentages:
Correctness: <points\>\#<reason\>
Hallucination: <points\>\#<reason\>
## Appendix CAdditional Results
Tables[9](https://arxiv.org/html/2608.19726#A3.T9)and[10](https://arxiv.org/html/2608.19726#A3.T10)are extensions of Tables[2](https://arxiv.org/html/2608.19726#S4.T2)and[3](https://arxiv.org/html/2608.19726#S4.T3)respectively, with additional rows for more existing models, joint variants at 16 hours, the curriculum learning ablation, and joint variants without LoRA adapters loaded\. Table[11](https://arxiv.org/html/2608.19726#A3.T11)reports the exact values plotted in Figure[3](https://arxiv.org/html/2608.19726#S4.F3)comparing projector\-only and joint training runs\.
Table 9:Generative 3D object classification results on ModelNet40 \(M40\.\), Objaverse \(Obj\.\), and OmniObject3D \(Omni\.\) under a zero\-shot setting\. Extension of Table[2](https://arxiv.org/html/2608.19726#S4.T2)\.Table 10:3D object captioning results on Objaverse\. Extension of Table[3](https://arxiv.org/html/2608.19726#S4.T3)\.Table 11:Generative zero\-shot classification accuracy against wall\-clock hours\. This is calculated as the average of the I and C variants of the ModelNet40 and Objaverse benchmarks, directly comparable to Figure[3](https://arxiv.org/html/2608.19726#S4.F3)\.
## Appendix DExample Generations
Table[12](https://arxiv.org/html/2608.19726#A4.T12)shows captions generated by P\-Llama8B and J\-Llama8B at 16 GPU hours\. They are often able to correctly identify shape, color, and object semantics, such as what the object is, though in some challenging settings \(e\.g\., low\-poly objects\) they fail to identify the object class correctly while still identifying textural details accurately\.
Table 12:Example generations on long captioning tasks\. The MLLMs are given the prompt “Caption this 3D model in detail\.” All prompts are independently provided to the MLLMs\. The answers shown in this table were generated by P\-Llama8B and J\-Llama8B at 16 hours of training, as well as the two PointLLM models\.Similar Articles
Query-based Cross-Modal Projector Bolstering Mamba Multimodal LLM
This paper proposes a query-based cross-modal projector that compresses visual tokens via cross-attention to improve Mamba-based multimodal LLMs, boosting both performance and throughput on vision-language benchmarks while eliminating the need for manual 2D scan order design.
Attending to Multimodal Generation One Token at a Time
This paper investigates token-level attention shifts in multimodal large language models during generation, revealing consistent patterns and proposing a simple test-time intervention that significantly improves task performance.
Alignment Is All You Need: Instruction-Free Training for General Audio-Language Models
This paper introduces an instruction-free alignment-only method for building large audio-language models by freezing the LLM and audio encoder, training only a lightweight projector on self-generated data, achieving competitive performance with less data than traditional multi-stage pipelines.
Large Language Models as Unified Multimodal Learners for Clinical Prediction
The paper proposes converting multimodal patient data (text, labs, vitals) into a single natural language sequence and fine-tuning LLMs for clinical prediction, achieving comparable or better performance than specialized fusion architectures across three tasks.
Learning Geometric Representations from Videos for Spatial Intelligent Multimodal Large Language Models
GeoVR enhances multimodal large language models with 3D awareness by restructuring their semantic latent space through geometric knowledge distillation from 3D foundation models using multiple geometric targets.