Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc

arXiv cs.CL Papers

Summary

This paper investigates fine-tuning small language models (0.6B-20B parameters) to generate syntactically correct MiniZinc models from natural language descriptions, proposing a cross-model error bootstrapping method that achieves up to 98% execution accuracy, though solution accuracy remains limited.

arXiv:2607.20456v1 Announce Type: new Abstract: Large language models excel at code generation for mainstream programming languages but struggle with rare, domain-specific languages such as MiniZinc, a constraint modeling language for combinatorial problems. We investigate whether targeted fine-tuning can teach small language models (0.6B to 20B parameters) to generate syntactically correct and semantically valid MiniZinc models from natural language problem descriptions. Our key finding is that syntax errors dominate failures when working with this domain specific language: the out-of-the-box execution accuracy of small language models such as Qwen3, LLaMa, Gemma, and GPT-OSS is near-zero. We propose a cross-model error bootstrapping approach that collects syntax errors from multiple LLM runs and leverage those to curate an error correction training dataset. This dataset allows us fine-tune small language models that consistently improves both direct code generation and chain-of-thought approaches across all model sizes. With self-reflection and ensembling, our approach achieves up to 98\% execution accuracy. In parallel, solution accuracy still remains at 35\%, indicating that while syntax is learnable, constraint reasoning remains a challenge. We contribute our fine-tuning pipeline, datasets, and models to opens-source for further research on text-to-model translation.
Original Article
View Cached Full Text

Cached at: 07/24/26, 05:16 AM

# Learn2Zinc: Fine-tuning Small Language Models for Text-to-Model Translation in MiniZinc
Source: [https://arxiv.org/html/2607.20456](https://arxiv.org/html/2607.20456)
Serdar Kadıoğlu1, 2and Karthik Uppuluri1 1AI Center of Excellence, Fidelity Investments 2Department of Computer Science, Brown University serdark@cs\.brown\.edu

###### Abstract

Large language models excel at code generation for mainstream programming languages but struggle with rare, domain\-specific languages such asMiniZinc, a constraint modeling language for combinatorial problems\. We investigate whether targeted fine\-tuning can teach small language models \(0\.6B to 20B parameters\) to generate syntactically correct and semantically validMiniZincmodels from natural language problem descriptions\. Our key finding is that syntax errors dominate failures when working with this domain specific language: the out\-of\-the\-box execution accuracy of small language models such as Qwen3, LLaMa, Gemma, and GPT\-OSS is near\-zero\. We propose a cross\-model error bootstrapping approach that collects syntax errors from multiple LLM runs and leverage those to curate an error correction training dataset\. This dataset allows us fine\-tune small language models that consistently improves both direct code generation and chain\-of\-thought approaches across all model sizes\. With self\-reflection and ensembling, our approach achieves up to 98% execution accuracy\. In parallel, solution accuracy still remains at 35%, indicating that while syntax is learnable, constraint reasoning remains a challenge\. We contribute our fine\-tuning pipeline, datasets, and models to opens\-source for further research on text\-to\-model translation\.

## 1Introduction

Optimization technology has achieved significant advancements, ranging from dramatic improvements in solver efficiency to the development of high\-level modeling languages to enhance usability\. Nevertheless, the fundamental decision\-making framework has remained unchanged for decades, adhering to the de factomodel\-and\-run strategy\. Within this status quo, users are required to manually convert problem descriptions into optimization models, which are subsequently processed by solvers to obtain solutions\.

Over the years, high\-level modeling languages such asMiniZinc\(Nethercoteet al\.[2007](https://arxiv.org/html/2607.20456#bib.bib1)\), CPMpy\(Guns[2019](https://arxiv.org/html/2607.20456#bib.bib2)\), and GAMS\(Bussieck and Meeraus[2004](https://arxiv.org/html/2607.20456#bib.bib3)\)have partially addressed the accessibility challenge by providing solver\-agnostic approaches that are powerful and flexible\. These modeling frameworks enable practitioners to focus on describing their problems without worrying about specific solution methods, making them especially useful for real\-world applications, where requirements often change over time\. However, the cognitive barrier of translating problem descriptions into formal constraint models persists\. This barrier is particularly acute, as domain experts who deeply understand their problem domain often lack the specialized knowledge required for formal modeling\. The resulting dependency on modeling experts creates operational bottlenecks and can lead to misinterpretation of domain\-specific requirements during the translation process\.

In parallel, Large Language Models \(LLMs\) have emerged as the communication medium with machines\(OpenAI and others[2024](https://arxiv.org/html/2607.20456#bib.bib5), Team and others[2025](https://arxiv.org/html/2607.20456#bib.bib6), DeepSeek\-AI and others[2025](https://arxiv.org/html/2607.20456#bib.bib7)\)\. While language models are powerful at interfacing with natural text, they struggle with the consistency and precision required in formal, declarative approaches, from basic type declarations to complex constraint relations\. They face significant challenges in handling the mathematical and logical reasoning needed for text\-to\-model translation\(Simchi\-Leviet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib26), Wasserkruget al\.[2025](https://arxiv.org/html/2607.20456#bib.bib27), Kadıoğluet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib24)\)111[https://skadio\.github\.io/text2model](https://skadio.github.io/text2model)\.

This current gap between understanding textual descriptions and turning them into problem formulations indicates that more work is needed for modeling assistants\.

A critical observation motivating this work is thatMiniZinc, as a domain\-specific language, has far less representation in typical LLM pretraining corpora compared to established languages such as Python, C\+\+, or Java\. For perspective, the GitHub topic page forMiniZinclists on the order of a hundred repositories, compared to millions for Python or JavaScript\. This limited presence means that LLMs have had minimal exposure toMiniZincsyntax during their pretraining\. Supporting this observation, four out of five models we tested; Qwen3, LLaMa, and Gemini, achieve 0\.0% execution accuracy onMiniZincgeneration, and the largest model \(GPT\-OSS\-20B\) reaches only 6\.0%\. This confirms thatMiniZincis clearly out\-of\-distribution for the current best small language models\. This raises a fundamental research question:Can targeted fine\-tuning teach small language models a domain\-specific programming language such asMiniZinc?\. This is exactly what we study in this paper\.

### 1\.1Our Contributions

Our contributions are as follows:

1. 1\.We present the first systematic study of fine\-tuning small language models from 0\.6B to 20B parameters forMiniZinccode generation\.
2. 2\.To make fine\-tuning possible, we introduce a cross\-model error bootstrapping approach that collects syntax errors from multiple LLM runs and leverage this to create a realistic error correction training dataset\.
3. 3\.We propose three fine\-tuning strategies of increasing sophistication and show that augmenting the training data with error correction examples consistently outperforms both direct code generation and chain\-of\-thought approaches across all model sizes\.
4. 4\.Our performance, when using an ensemble of our fine\-tuned models,achieves 98% execution accuracy, up from 0\.0\-6\.0% out\-of\-the\-box accuracy, effectively solving theMiniZincsyntax problem\. At the same time, we identify constraint reasoning as the remaining bottleneck with only 35% solution accuracy with a detailed analysis on error modes\.
5. 5\.

## 2Background

Let us briefly review the constraint modeling language used in this study,MiniZinc, the dataset used for benchmarking,Text2Zinc, and our LLM copilot approaches,Text2Model\.

### 2\.1MiniZinc

MiniZinc\(Nethercoteet al\.[2007](https://arxiv.org/html/2607.20456#bib.bib1)\)is a high\-level constraint modeling language that supports both discrete and continuous optimization and satisfaction problems\. Its solver\-agnostic design allows communication with various solver backends, including Constraint Programming \(CP\), \(Mixed\) Integer Programming \(MIP\), and Boolean Satisfiability/Lazy Clause Generation \(SAT\)\. This flexibility is achieved through compilation toFlatZinc, an intermediate language that interfaces with different solvers, allowing the sameMiniZincmodel to be used across multiple backends without code modifications\.

A key feature ofMiniZincis its use of global constraints, which significantly simplify the modeling process\. For example, theall\_differentconstraint specifies that a set of variables must take distinct values, replacing numerous pairwise inequality constraints\.

TheMiniZinclanguage structure consists of four main components: decision variables, constraints, parameters, and an objective function \(for optimization\) or a satisfaction goal\.MiniZincalso separates models \(\.mznfiles\) from data instances \(\.dznfiles\), allowing a single model to be reused across multiple problem instances\.

This paper builds on two previous complementary works:Text2Zincdataset to establish a common benchmark for this task, andText2Modelcopilots to establish a baseline performance of various LLM approaches\.

### 2\.2Text2ZincDataset

Text2Zinc\(Singirikondaet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib41)\)introduces a cross\-domain dataset for modeling optimization and satisfaction problems inMiniZinc\.It is the first dataset in this line of research that covers both satisfaction and optimization problems in a solver\- and paradigm\-agnostic language\.The dataset brings together 1,775 problems from multiple sources, includingNlp4lp\(AhmadiTeshniziet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib19)\),Hakank,ComplexOr\(Xiaoet al\.[2023](https://arxiv.org/html/2607.20456#bib.bib17)\),LpWp\(Ramamonjisonet al\.[2022](https://arxiv.org/html/2607.20456#bib.bib16)\),CspLib, and problems from Cardinal Operations covering bothMamoandNl4Opt\(Kadıoğluet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib24)\)collections \(see Table[1](https://arxiv.org/html/2607.20456#S3.T1)for a full breakdown by category\)\. Of these, 110 are fully verified with manually writtenMiniZincmodels, complete metadata, and validated solutions\. For the remaining problems, which originate fromIndustryOr,Mamo, andNl4Opt, ground\-truth objective values are available through the original sources for verification purposes\. The dataset providesi​s​\_​o​p​t​i​m​i​z​a​t​i​o​nis\\\_optimization,i​s​\_​s​a​t​i​s​f​a​c​t​i​o​nis\\\_satisfaction,h​a​s​\_​v​e​r​i​f​i​e​d​\_​o​b​jhas\\\_verified\\\_obj,h​a​s​\_​v​e​r​i​f​i​e​d​\_​m​z​nhas\\\_verified\\\_mzn,h​a​s​\_​d​z​nhas\\\_dznto distinguish among these properties\.

### 2\.3Text2ModelCopilots

Text2Model\(Kadıoğluet al\.[2026](https://arxiv.org/html/2607.20456#bib.bib42)\)introduces a suite of copilots for text\-to\-model translation using frontier LLMs\. It evaluates several strategies of varying complexity on theText2Zincdataset, including zero\-shot prompting, chain\-of\-thought reasoning, knowledge\-graph representations, grammar\-based syntax encoding, and agentic approaches\. A key observation is that even frontier LLMs with sophisticated prompting strategies are notyeta push\-button technology for combinatorial modeling\.

## 3Learn2Zinc: Fine\-Tuning Small Language Models

The important insight that led to this work is as follows\. Running severalText2Modelcopilots on hundreds ofText2Zincinstances generate a considerable amountMiniZincmodels, even though theMiniZincmodels are not manually written\. While generation is straightforward, the key advantage stems from theverification loop: the output can be verified with known objective value for optimization problems, and feasibility can be asserted for satisfaction problems\. This yields a large pool of verified⟨t​e​x​t,m​o​d​e​l⟩\\langle text,model\\ranglepairs that makes fine\-tuning possible\.

In this paper, we go beyond dependency on large frontier models for generating constraint models and investigate whether targeted fine\-tuning can teachsmall language models\(0\.6B–20B parameters\) to generateMiniZinccode\. Crucially, this effort requires designing a fine\-tuning dataset\. Our fine\-tuning dataset starts from the verifiedMiniZincsolutions inText2Zinc\. In addition, we draw samples leveraging theOr\-Instructdataset\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\)555We are indebted to the creators of theOr\-Instructdataset for their valuable contribution to the community\.\. TheOr\-Instructdataset contains optimization problems withCoptmodels written in Python\. Given these problem descriptions, theirCoptmodel, and verified objective value, we use GPT\-5\.2 to generate a translation fromCopttoMiniZinc\. We add the result to our fine\-tuning dataset if and only ifMiniZincoutput matches the known objective\.

As shown in Table[1](https://arxiv.org/html/2607.20456#S3.T1), whenMiniZincmodels forText2ZincandOr\-Instructare combined, they together cover a set of2,208unique problems\. Importantly, instances might be associated with alternativeMiniZincmodels generated by differentText2Modelcopilot strategy\. Overall, this yields8,014instruction\-tuning⟨t​e​x​t,m​o​d​e​l⟩\\langle text,model\\ranglepairs for fine\-tuning small language models\.

DatasetSourceInstancesPercentageOrlm\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\)Or\-Instruct\-3K1,35161\.2%Text2Zinc\(Singirikondaet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib41)\)Mamo60427\.4%Nl4Opt2009\.1%Nlp4Lp301\.4%Hakank100\.5%ComplexOr50\.2%LpWp50\.2%CspLib30\.1%Total2,208100%Table 1:Distribution of unique problem instances by source\. Problems fromText2Zincwere verified against known objectives usingText2Modelcopilot strategies\.Or\-Instructproblems were translated fromCoptmodels written in Python toMiniZincvia GPT\-5\.2\.
## 4Learn2Zinc: Fine\-tuning Methodology

Our fine\-tuning methodology consists of a combination of different small language models at varying parameter sizes and different fine\-tuning datasets\. We consider both ageneration fine\-tuninganderror\-correction fine\-tuningto teach SLMsMiniZincsyntax \(§[5](https://arxiv.org/html/2607.20456#S5)\)\.

### 4\.1Small Language Models \(Slms\)

We consider four different families of small languages models \(SLMs\): Qwen, LLaMa, Gemma, and GPT\. In particular, we fine\-tune five different variants to cover different parameter sizes: Qwen3\-0\.6B \(the smallest to test the lower bound of capability\), LLaMA\-3\.2\-1B and LLaMA\-3\.2\-3B, Gemma\-2\-9B, and GPT\-OSS\-20B \(the largest number of parameters\)\.

### 4\.2Datasets & Setup

We consider three different fine\-tuning instruction datasets to consider a baseline, a chain\-of\-thought, and a mixed strategy\.

#### Learn2Zinc\-Base\.

We obtain 8,014 instances with⟨t​e​x​t,m​o​d​e​l⟩\\langle text,model\\ranglepairs as described in §[4](https://arxiv.org/html/2607.20456#S4)\.

#### Learn2Zinc\-CoT\(Chain\-of\-Thought\)\.

We take the baseline dataset with⟨t​e​x​t,m​o​d​e​l⟩\\langle text,model\\ranglepairs, and prompt a reasoning model, GPT4o\. to generate reasoning chain from the problem text to the constraint model and identify variables, parameters, constraints, objective\. Accordingly, this dataset also has 8,014 pairs of the form⟨t​e​x​t,r​e​a​s​o​n​i​n​g,m​o​d​e​l⟩\\langle text,reasoning,model\\rangle\.

#### Learn2Zinc\-Base\+CoT\.

We combine bothLearn2Zinc\-BaseandLearn2Zinc\-CoTexamples leading to 16,028 pairs\.

In Appendix[A](https://arxiv.org/html/2607.20456#A1), we present examples of our baseline and CoT pairs\. Notice how the CoT version introduces the reasoning chain as an intermediary in between the problem description and the constraint model\.

In our experiments, we fine\-tune five different SLMs on three different fine\-tuning datasets\. The details of training hyperparameters are given in Appendix\-Table[10](https://arxiv.org/html/2607.20456#A2.T10)\. We employ Low\-Rank Adaptation \(LoRA\) with 8\-bit quantization for all models, except 4\-bit for GPT\-OSS\-20B, to fit on a single A100 GPU\.

We choose LoRA over full fine\-tuning for two reasons\. First, our fine\-tuning datasets are relatively small, with the base dataset comprising only 8,014 instances\. At this scale, full fine\-tuning is prone to overfitting and offers limited gains over parameter\-efficient alternatives\. Second, LoRA significantly reduces memory and compute requirements, allowing us to fine\-tune models ranging from 0\.6B to 20B parameters on a single GPU, where full fine\-tuning would not be feasible\. Among parameter\-efficient fine\-tuning \(PEFT\) approaches, LoRA is well\-established and widely adopted, offering a strong balance between task adaptation and training efficiency, making it a suitable choice for our experimental setting\.

ModelStrategyExec Acc \(%\)Sol Acc \(%\)Qwen3\-0\.6BOriginal\-8bit0\.00\.0Learn2Zinc\-Base51\.09\.0Learn2Zinc\-CoT44\.010\.0Learn2Zinc\-Base\+CoT51\.012\.0Learn2Zinc\-Augmented64\.013\.0LLaMA\-3\.2\-1BOriginal\-8bit0\.00\.0Learn2Zinc\-Base44\.04\.0Learn2Zinc\-CoT22\.01\.0Learn2Zinc\-Base\+CoT46\.04\.0Learn2Zinc\-Augmented57\.08\.0LLaMA\-3\.2\-3BOriginal\-8bit0\.00\.0Learn2Zinc\-Base62\.010\.0Learn2Zinc\-CoT47\.09\.0Learn2Zinc\-Base\+CoT58\.012\.0Learn2Zinc\-Augmented70\.017\.0Gemma\-2\-9BOriginal\-8bit0\.00\.0Learn2Zinc\-Base74\.022\.0Learn2Zinc\-CoT49\.015\.0Learn2Zinc\-Base\+CoT70\.021\.0Learn2Zinc\-Augmented72\.022\.0GPT\-OSS\-20BOriginal\-4bit6\.05\.0Learn2Zinc\-Base66\.027\.0Learn2Zinc\-CoT52\.021\.0Learn2Zinc\-Base\+CoT63\.027\.0Learn2Zinc\-Augmented76\.032\.0GPT\-5\.2†Agentic \+ Code\(Kadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\)\)86\.057\.0Table 2:Execution and Solution Accuracy \(%\) across models and fine\-tuning strategies\. Best results per model are highlighted\.†GPT\-5\.2 results are fromKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\)and shown for reference\.

### 4\.3Evaluation Metrics

We use Execution Accuracy \(the constraint model compiles and runs\) and Solution Accuracy \(the objective value found by solving the constraint model matches the ground truth\) as used previously in several studies, e\.g\.,Kadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\)\.

### 4\.4Numerical Results

For benchmarking, we useText2Zinc\-IndustryOr\(100 problems\) as our test set\. It is important to note theses problems arenot included in the fine\-tuning\. We useIndustryOrsince it is the most challenging benchmark, as reported inKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\), Huanget al\.\([2024](https://arxiv.org/html/2607.20456#bib.bib23)\), Lianget al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib40)\), which shows 86% execution accuracy and 57% solution accuracy with an Agentic copilot using the frontier GPT\-5\.2 model\. As the underlying solver, we use HiGHS in our experiments, with a 120\-second timeout per problem\.

Table[2](https://arxiv.org/html/2607.20456#S4.T2)presents execution and solution accuracy for each model and fine\-tuning strategy\. Please omit the augmented row in the table for now as we discuss this later in §[5](https://arxiv.org/html/2607.20456#S5)\. Figure[1](https://arxiv.org/html/2607.20456#S4.F1)visualizes the same results to reveal emerging patterns\.

Several patterns emerge from results in Table[2](https://arxiv.org/html/2607.20456#S4.T2)\. First of all,original SLMs achieve near\-zero accuracy\. All models score 0%, except for GPT\-OSS\-20B with 6\.0% execution and 5\.0% solution accuracy\. This confirms the main motivation of this paper:MiniZincis clearly out\-of\-distribution for current best SLMs\.With our fine\-tuning, Gemma\-2\-9B achieves a remarkable boost from its original 0% execution accuracy to 74%\. As a comparison between SLMs and frontier models, according toKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\),GPT\-5\.2 performs 86% execution and 57% solution accurac on the same test setto generateMiniZincmodels, as reported at the bottom row of Table[2](https://arxiv.org/html/2607.20456#S4.T2)\.

As Figure[1](https://arxiv.org/html/2607.20456#S4.F1)makes it clear, interestingly,Learn2Zinc\-CoTunderperformsLearn2Zinc\-Baseacross all models\. This finding contrasts with several existing results where chain\-of\-thought often helpsKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\), Singirikondaet al\.\([2025](https://arxiv.org/html/2607.20456#bib.bib41)\), hinting that, unlike frontier models, SLMs cannot take advantage of reasoning traces\. A possible interpretation is, reasoning traces cannot compensate for missing syntax when the small model has not been trained forMiniZincat scale\. Similarly, simply mixingLearn2Zinc\-BaseandLearn2Zinc\-CoTexamples provides no consistent improvement overLearn2Zinc\-Basealone\. Based on these findings, we therefore pursue a dedicated approach next to teach SLMsMiniZincsyntax\.

Qwen3\-0\.6BLLaMA\-1BLLaMA\-3BGemma\-9BGPT\-OSS\-20B02020404060608080100100Execution Accuracy \(%\)Learn2Zinc\-BaseLearn2Zinc\-CoTLearn2Zinc\-Base\+CoTLearn2Zinc\-AugmentedFigure 1:Execution accuracy \(%\) across fine\-tuning strategies and model sizes\.Learn2Zinc\-Augmented\(green\) consistently achieves the highest execution accuracy across all models, whileLearn2Zinc\-CoT\(blue\) underperforms all other strategies\.

## 5Teaching SlmsMiniZincSyntax

Our initial fine\-tuning experiments with five SLMs and three fine\-tuning strategies reveals a consistent pattern:SLM outputs are riddled with syntax errors\. SLMs, even with our careful fine\-tuning, cannot produce correctMiniZinccode where the best execution accuracy is still at 76%\. Before we could even evaluate whether SLMs capture correct combinatorial semantics, most outputs fail to captureMiniZincsyntax and compile\. Execution accuracy, even before solution accuracy, remains an immediate bottleneck\.

This observation shifts our focus\. If we could potentially improve execution rates, models would have more opportunity to produce correct solutions\. The question then becomes:how can we teach SLMs to avoid and correct their syntax mistakes?So far, we approached fined\-tuning for text\-to\-model translation asgeneration fine\-tuningand we now shift our focus toerror\-correction fine\-tuning\.

For this purpose, we design an augmented training strategy thatexplicitly teaches syntax correction\. The key insight is that models need exposure to common errorsand their fixes, not just correct outputs\.

Rule IDDifficultyDescriptionExampleE1\_drop\_semicolonEasyRemove trailing semicolon from a statementvar int: x = 5;→\\rightarrowvar int: x = 5E2\_drop\_commaEasyRemove a comma from a list, array, or parameter list\[1, 2, 3\]→\\rightarrow\[1 2, 3\]E3\_2d\_array\_openEasyReplace \[\| with \[ in 2D array literal\[\| 1, 2 \| 3, 4 \|\]→\\rightarrow\[ 1, 2 \| 3, 4 \|\]E4\_2d\_array\_closeEasyReplace \|\] with \] in 2D array literal\[\| 1, 2 \| 3, 4 \|\]→\\rightarrow\[\| 1, 2 \| 3, 4 \]E5\_drop\_row\_pipeEasyRemove \| row separator in 2D array\[\| 1, 2 \| 3, 4 \|\]→\\rightarrow\[\| 1, 2 3, 4 \|\]E6\_drop\_close\_parenEasyRemove a closing parenthesissum\(x\)→\\rightarrowsum\(xE7\_drop\_close\_bracketEasyRemove a closing bracketarray\[1\.\.N\]→\\rightarrowarray\[1\.\.NM1\_drop\_varMediumRemove "var" keyword from variable declarationvar int: x→\\rightarrowint: xM2\_drop\_ofMediumRemove "of" keyword from type expressionset of int→\\rightarrowset intM3\_solve\_keywordMediumReplace "maximize"/"minimize" with invalid "max"/"min"solve maximize obj→\\rightarrowsolve max objM4\_capitalize\_keywordMediumCapitalize aMiniZinckeyword \(case\-sensitive error\)constraint x \> 0→\\rightarrowConstraint x \> 0M5\_split\_endifMediumSplit "endif" into "end if" or "elseif" into "else if"endif→\\rightarrowend ifM6\_drop\_gen\_inMediumRemove "in" from generator expressionforall\(i in 1\.\.N\)→\\rightarrowforall\(i 1\.\.N\)H1\_drop\_thenHardRemove "then" from if\-then\-else expressionif x \> 0 then y→\\rightarrowif x \> 0 yH2\_drop\_let\_inHardRemove "in" keyword after let block closing bracelet \{ var int: t \} in t→\\rightarrowlet \{ var int: t \} tH3\_swap\_type\_identHardSwap type and identifier in declarationvar int: x→\\rightarrowx: var intH4\_triple\_dot\_rangeHardReplace "\.\." range operator with "…" \(invalid\)1\.\.N→\\rightarrow1…NH5\_wrong\_logic\_opHardReplace/\\with&&or\\/with\|\|\(invalid C\-style operators\)x /\\ y→\\rightarrowx && yH6\_extra\_semicolonHardInsert extra semicolon in the middle of a constraintconstraint x \+ y = z→\\rightarrowconstraint x \+; y = zH7\_drop\_close\_braceHardRemove a closing brace \}\{1, 2, 3\}→\\rightarrow\{1, 2, 3Table 3:Complete taxonomy ofMiniZincsyntax corruption rules derived from the BNF grammar\.### 5\.1Error Correction Dataset

For error correction fine\-tuning, we draw data samples from two different strategies; synthetic corruption and cross\-model error bootstrapping\.

#### Synthetic Corruptions\.

We leverage the Backus\-Naur Form \(BNF\) grammar ofMiniZinc666We thank Guido Tack for sharing this grammar with us\.to derive 20 corruption rules, each grounded in a specific grammar production\. These rules are organized into three difficulty levels: easy \(7 rules targeting delimiters and punctuation, e\.g\., missing semicolons, dropped commas, unbalanced brackets\), medium \(6 rules targeting keywords and types, e\.g\., invalid solver directives, droppedvarorofkeywords, split compound tokens likeendif\), and hard \(7 rules targeting structural elements, e\.g\., swapped declaration order, invalid operators, misplaced semicolons\)\.

Table[3](https://arxiv.org/html/2607.20456#S5.T3)lists the complete taxonomy of grammar\-based corruption rules\. To generate the training data, we sample instances from the baseline dataset and apply corruptions at controlled difficulty with 30% easy, 30% medium, 30% hard, and 10% identity\. Within each difficulty we sample the rule to apply proportionally with our empirical observations\. In case of identity, the model needs to recognize that no fix is necessary\. Each corrupted model is verified using theMiniZinccompiler to ensure the gold code passes and the corrupted code fails syntax checking; pairs that do not satisfy both conditions are discarded\. This process yields4,452 verified instanceswith⟨t​e​x​t,m​o​d​e​l,c​o​r​r​u​p​t​\_​m​o​d​e​l⟩\\langle text,model,corrupt\\\_model\\rangletriples\. Figure[2](https://arxiv.org/html/2607.20456#S5.F2)shows the distribution of rules applied in the resulting dataset color coded by their difficulty\.

![Refer to caption](https://arxiv.org/html/2607.20456v1/fig_rules_by_difficulty.png)Figure 2:Top\-15MiniZincgrammar corruption rules color coded by difficulty\.
#### Cross\-Model Error Bootstrapping\.

Our probabilistic synthetic error generation fromMiniZincgrammar may not capture actual failure modes of LLMs at test time\. We address this by collecting errors and their fixed version as follows:

1. 1\.We run all five SLMs across all three fine\-tuning strategies \(base, CoT, mixed\) on the base dataset with multiple sampling temperatures to increase variety of error cases\.
2. 2\.When we obtain an execution failure, for each⟨t​e​x​t,c​o​r​r​u​p​t​\_​m​o​d​e​l,e​r​r​o​r​\_​m​e​s​s​a​g​e⟩\\langle text,corrupt\\\_model,error\\\_message\\rangle, we employ GPT\-5\.2 to attempt a fix\. A natural alternative would be to pair the failed output directly with a correct generation from the same model\. However, we found that correct and incorrect generations from the same SLM often differ substantially in structure and variable naming, meaning the pairing would not represent a targeted fix but rather a complete rewrite\. Such examples would poorly represent the error correction task\. Instead, we prompt GPT\-5\.2 with the corrupted model and its error message, instructing it to make only the minimal changes necessary to fix the error while preserving the original code structure\. This produces training pairs where the correction is localized to the actual fault\. 1. \(a\)If the fix succeeds, i\.e, the updated model compilesandproduces the correct output, we add⟨t​e​x​t,c​o​r​r​u​p​t​\_​m​o​d​e​l,c​o​r​r​e​c​t​\_​m​o​d​e​l⟩\\langle text,corrupt\\\_model,correct\\\_model\\rangleto our fine\-tuning dataset\. This process yields2,286verified instances and is aimed at teaching SLMS error correction\. Table[4](https://arxiv.org/html/2607.20456#S5.T4)shows the distribution of error\-correction examples over SLMs varied by temperature\. In total, the cross\-model error bootstrapping process creates a diverse error corpus\. As expected, smaller models \(Qwen3\-0\.6B and LLaMA\-3\.2\-1B\) contribute the most correction examples\. Fine\-tuning SLMs with this dataset has an interesting property: small models \(e\.g\., Qwen3\-0\.6B\) learn from large model \(e\.g\., GPT\-OSS\-20B\) mistakes, and vice versa\. Higher sampling temperatures also produce more errors with temperature 0\.8 accounting for 38\.1% of corrections vs\. 29\.4% at temperature 0\.2\. Table[5](https://arxiv.org/html/2607.20456#S5.T5)shows the distribution of errors in this cross\-model dataset by the fine\-tuning strategy\. Notice that,Learn2Zinc\-CoTtraining produces the most errors during bootstrapping \(43\.9%\), suggesting again that reasoning traces without syntax knowledge may actually introduce more mistakes\.
3. 3\.Alternatively, when we obtain executionandsolution success, we collect a new positive generation example,⟨t​e​x​t,c​o​r​r​e​c​t​\_​m​o​d​e​l⟩\\langle text,correct\\\_model\\rangleto our fine\-tuning dataset\. This process yields8,911 verified instances, i\.e\., slightly extends our base dataset from 8,014, and is aimed at teaching SLMs constraint model generation\.

ModelT=0\.2T=0\.5T=0\.8Total%LLaMA\-3\.2\-1B23026228677834\.0Qwen3\-0\.6B21725627474732\.7LLaMA\-3\.2\-3B13912619245720\.0Gemma\-2\-9B8510111830413\.3Total6717458702,286100\.0%29\.432\.638\.1Table 4:Distribution of correction pairs by error\-producing SLM and sampling temperatures\.StrategyErrorsPercentageLearn2Zinc\-Base64028\.0%Learn2Zinc\-CoT1,00343\.9%Learn2Zinc\-Base\+CoT64328\.1%Total2,286100%Table 5:Errors by training strategy during bootstrapping\.Task TypeCountError\-Correction tasks \(total\)6,738\- Synthetic corruptions4,452\- Cross\-model corrections2,286Generation task8,911Total15,649Table 6:Breakdown of the augmented dataset by task type\.Overall, as shown in Table[6](https://arxiv.org/html/2607.20456#S5.T6), when we combine cross\-model error correction bootstrapping instances \(6,728\) and generation task instances together \(8,911\), we obtain15,649 instancesfor fine\-tuning\.

### 5\.2Learn2Zinc\-AugmentedFine\-Tuning

Given the the augmented fine\-tuning dataset in Table[6](https://arxiv.org/html/2607.20456#S5.T6), which hosts 15,649 instances in combination, we conduct the same fine\-tuning protocol on all five SLMs\. This produces ourLearn2Zinc\-Augmentedfine\-tuned model families\. Notice that compared to our previous fine\-tuning in §[4](https://arxiv.org/html/2607.20456#S4), this augmented fine\-tuning has two simultaneous objectives; a generation task and an error\-correction task\.

As previously reported in Table[2](https://arxiv.org/html/2607.20456#S4.T2), theLearn2Zinc\-Augmentedstrategy achieves the best execution accuracy on all models, except one, and more importantly, achieves the best solution accuracy among all approaches\. When analyzing SLMs individually, Qwen3\-0\.6B jumps from 0\.0% execution to 51% with our initial fine\-tuning, and then to 65% accuracy with our augmented fine\-tuning\. Similarly, LLaMa\-3\.2\-1B goes from 0\.0% to 46% and then to 57%\. Gemma\-2\-9B goes from 0\.0% to 74% and 72%\. GPT\-OSS\-20B goes from 6\.0% to 66% and then to 76%\.Our best zero\-shot results achieves 76% using augmented GPT\-OSS\-20B\. This validates thattraining for error correction directly addresses the syntax bottleneck\.

## 6Learn2Zincwith Self\-Reflection & Ensemble

So far, we only considered a zero\-shot strategy, even though we fined\-tuned for error correction\. As such, we next utilize aself\-reflection loopandensembleon top of theLearn2Zinc\-Augmentedstrategies\. Since other fine\-tuning strategies have never seen error correction examples, we only consider results ofLearn2Zinc\-Augmentedstrategy for self\-reflection and ensembling\.

### 6\.1Self\-Reflection Loop

We run augmented fine\-tuned models up to five attempts to produce executable code\. On each failed attempt, the model receives its broken code and the compiler error message, then tries again, until hitting the limit\.

As shown in Table[7](https://arxiv.org/html/2607.20456#S6.T7), adding reflection improves execution accuracy considerably and slightly improves solution accuracy\. Notice that reflections reach a running model, on average,in less than two trials, with larger SLMs requiring fewer attempts\.

Single PassSelf\-Reflection@5ModelExec \(%\)Sol \(%\)Exec \(%\)Sol \(%\)Avg AttemptsQwen3\-0\.6B\-Augmented64\.013\.072\.013\.02\.22LLaMA\-3\.2\-1B\-Augmented57\.08\.071\.08\.02\.35LLaMA\-3\.2\-3B\-Augmented70\.017\.080\.017\.01\.90Gemma\-2\-9B\-Augmented72\.022\.084\.024\.01\.79GPT\-OSS\-20B\-Augmented76\.032\.089\.034\.01\.62Table 7:Learn2Zinc\-Augmented Slms: Single\-pass vs\. Self\-reflection with five attempts\.Execution accuracy improves substantially with retries, confirming thatLearn2Zinc\-Augmentedtraining teaches models to fix their own mistakes at inference time\.Our augmented GPT\-OSS\-20B reaches 89% execution accuracy, up from 76% on the initial fine\-tuning, and significantly up from its original 6\.0%\.In other words, our augmented fine\-tuning with self\-reflection enables GPT\-OSS achieve an on par execution accuracy \(89%\) compared to its frontier variant, GPT\-5\.2 \(86%\)\. This is one of our main contributions\.

However, solution accuracy remains flat\. SLMs that fail to solve a problem on their first executable attempt do not succeed later\. This suggests a separation between syntax errors \(fixable through retry\) and semantic errors \(not fixable due to SLM capacity\)\. Once the code compiles, the model has already committed to a constraint formulation\. If that formulation is wrong, producing further executable code does not help\.

### 6\.2Ensemble

Finally, as an alternative to self\-reflection and instead of multiple attempts usingthe same SLM, we evaluate a top\-down ensemble \(largest SLM first\) to exploit the variation across SLMs\. In our ensemble, models are tried in descending order of capability: GPT\-OSS\-20B→\\rightarrowGemma\-2\-9B→\\rightarrowLLaMA\-3\.2\-3B→\\rightarrowLLaMA\-3\.2\-1B→\\rightarrowQwen3\-0\.6B\. If a model fails to produce executable code after retries, the next model is attempted\. We consider a top\-down strategy to increase the chance of getting an executable model where the solution accuracy remains highest \(stronger models\)\. Conversely, if we build this ensemble bottom\-up \(smallest SLM first\) we risk obtaining an executing model when the solution reasoning capacity is lowest\.

Table[8](https://arxiv.org/html/2607.20456#S6.T8)presents the main takeaways from our workLearn2Zinc\. We start with the original GPT\-OSS\-20B, the largest of our SLMs, which has an extremely poor performance on generatingMiniZincmodels with only 6% execution and 5% solution accuracy\. Next, our fine\-tuning, the initial base strategy,Learn2Zinc\-Base, detailed in §[4](https://arxiv.org/html/2607.20456#S4), and its augmented version,Learn2Zinc\-Augmentedfor error\-correction detailed in §[5](https://arxiv.org/html/2607.20456#S5)steadily improveMiniZinccoding abilities of GPT\-OSS with up to 76% execution accuracy\. The solution accuracy also jumps from 5% to 32%\. Our augmented version,Learn2Zinc\-Augmented, has the additional property to serve well in self\-reflection loops to fix its coding errors, which it is exactly trained for, that leads to another boost in execution accuracy to 89%\.Finally, our ensemble of fined\-tuned SLMS achieves a remarkable 98% execution accuracy closing the syntactic aspect of the text\-to\-model translation task\. As a reference, the frontier GPT\-5\.2 achieves only 86% execution accuracy forMiniZinccode with the best copilot,Agentic \+ Code, fromKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\)\. We show that it is possible to teach SLMs domain specific languages for text\-to\-model translation via fine\-tuning\.

Overall, this is the central finding of our approach: withLearn2Zinc\-Augmentedfine\-tuning, self\-reflection, and SLM ensembles, the execution problem is effectively solved forMiniZinc\.Our method allows pushing bottleneck to shift entirely to semantic reasoning about constraint modeling\.

Beyond the syntax, what remains is the gap in solution accuracy for formal reasoning\. Our SLM fine\-tuning saturates around 35% on solution accuracy\. Even the frontier models such as GPT\-5\.2 struggles with this at 57%\. The best known solution onIndustrORcomes fromLianget al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib40)\)with 65% using Gemini 3 Pro for generatingGurobimodels in Python, as we discuss next\.

ModelApproachExecution \(%\)Solution \(%\)GPT\-OSS\-20BOriginal\-4bit6\.05\.0GPT\-OSS\-20BLearn2Zinc\-Base66\.027\.0GPT\-OSS\-20BLearn2Zinc\-Augmented76\.032\.0GPT\-OSS\-20BLearn2Zinc\-Aug\+Self\-Reflection89\.034\.0Our fined\-tuned SLMsLearn2Zinc\-Ensemble98\.035\.0GPT\-5\.2Agentic \+ Code\(Kadıoğluet al\.[2026](https://arxiv.org/html/2607.20456#bib.bib42)\)86\.057\.0Table 8:Learn2Zincoverall comparison with different strategies\.

## 7Learn2Zincvs\. Prior Work

We compareLearn2Zincwith prior work on the sameIndustryOrbenchmark fromOrlm\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\)andLean\-Lllm\-Opt\(Lianget al\.[2026](https://arxiv.org/html/2607.20456#bib.bib40)\)\. Before numerical comparisons, let us first highlight that such a comparison is subject to several caveats and should be treatedonly directionally\.

First, target constraint languages differ\.Orlm\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\)generates Python code for theCoptsolver whereasLean\-Lllm\-Opt\(Lianget al\.[2026](https://arxiv.org/html/2607.20456#bib.bib40)\)generates Python code for theGurobisolver\.

Python is a general\-purpose programming language and is abundant in the LLM pretraining data\. As such, LLMs produce syntactically valid code easily for Python\. Contrarily, our approach generatesMiniZinc, a domain\-specific modeling language, and as shown in our experiments, remains out\-of\-distribution, especially for SLMs\. An important contribution ofLearn2Zincis to address syntax challenge which is taken for granted in Pythonic approaches\.

Second, evaluation protocols differ\.Orlmreports Pass@k by sampling k independent outputs and counting success ifanyone is correct\. Our evaluation protocol reflects an agentic loop: we iterate until execution succeeds or a budget is exhausted, while sharing error messages in between, and counting the total number of LLM calls in average termination\. These results are not directly comparable, but we report both for context\.

Third, the underlying LLMs differ\.Orlmfine\-tunes over LLaMA\-3\-8B, similar but not exactly identical to our case, whileLean\-Lllm\-Optis based on GPT\-4\.1 and GPT\-OSS\-20B\. For a complete picture, we also include results on GPT and Gemini for generating Gurobi\-Python from\(Lianget al\.[2026](https://arxiv.org/html/2607.20456#bib.bib40)\), as\-is, as well as results on GPT for generatingMiniZincfrom\(Kadıoğluet al\.[2026](https://arxiv.org/html/2607.20456#bib.bib42)\)as\-is\.

Table[9](https://arxiv.org/html/2607.20456#S7.T9)shows results onIndustryOrbenchmark forOrlm\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\),Lean\-Lllm\-Opt\(Lianget al\.[2026](https://arxiv.org/html/2607.20456#bib.bib40)\),Text2Model\(Kadıoğluet al\.[2026](https://arxiv.org/html/2607.20456#bib.bib42)\), andLearn2Zinc\. Our best fine\-tuned GPT\-OSS\-20B achieves 34% onIndustryOr, below the Pythonic methods\. This gap reflects two factors\. First,MiniZincsyntax remains harder than Python even thoughLearn2Zinc\-Augmentedcan address it\. Second, our SLM are lighter than the frontier models used by other methods\. TheText2Modelresult \(57% with GPT\-5\.2\) shows that stronger models withMiniZinccan approach Pythonic performance\.

MethodModelLLM CallsIndustryOr\(%\)Fine\-tuning approach forCoptin PythonORLM \(Pass@1\)LLaMA\-3\-8B138\.0ORLM \(Pass@8\)LLaMA\-3\-8B849\.0Agentic framework forGurobiin PythonLean\-Lllm\-Opt†GPT\-OSS\-20B5\+59\.0Lean\-Lllm\-Opt†GPT\-4\.15\+65\.0Direct prompting forGurobiin Python†–GPT\-OSS\-20B142\.0–GPT\-4\.1154\.0–GPT\-5158\.0–GPT\-5\.2156\.0–Gemini 3 Pro165\.0Text2Modelcopilots forMiniZincText2Model: Agentic \+ CodeGPT\-OSS\-20B15\.0Text2Model: Agentic \+ CodeGPT\-5\.2149\.0Text2Model: Agentic \+ CodeGPT\-5\.2557\.0Fine\-tuningMiniZinc\(this paper\)Learn2Zinc\-AugmentedGPT\-OSS\-20B132\.0Learn2Zinc\-Augmented\+Self\-ReflectionGPT\-OSS\-20B1\.6234\.0Table 9:IndustryOrcomparisons acrossOrlm, generatingCoptmodels in Python,Lean\-Lllm\-OptgeneratingGurobimodels in Python,Text2ModelandLearn2ZincgeneratingMiniZincacross different LLMs and SLMs\. Results marked with†\\daggerare fromLianget al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib40)\)\.The value ofLearn2Zincis complementary and methodological: we produce solver\- and paradigm\- agnosticMiniZincmodels that can solve optimizationandsatisfaction problems, whereasOrlmis tied toCoptandLean\-Lllm\-Optis tied toGurobi\. We maintain solver and paradigm flexibility while addressing SLMs’s disadvantage on domain\-specific languages viaLearn2Zincfine\-tuning\.

## 8Discussions & Error Analysis

Our immediate obsersation is thatout\-of\-the\-box SLMs achieve near 0% performance\. Three factors combine to makeMiniZincout\-of\-distribution for current LLMs\. First,MiniZinchas minimal GitHub presence compared to mainstream languages\. Second, it uses a declarative paradigm unlike the imperative code that makes up most pretraining data\. Third, no syntactically similar language exists for transfer learning\. The combination of rare syntax and unfamiliar paradigm explains the complete failure of unfinetuned models\.

We also notice thatCoT is not effective when fine\-tuning SLMsasLearn2Zinc\-CoTunderperforms\. When prompting an LLM, CoT often helps because models already know the target language and they benefit from structured reasoning but fine\-tuning is different\.Learn2Zinc\-CoTrequires the SLM to reason about validMiniZincconstructs it has seldom or never seen\. SLMs might not reason about syntax that they have not acquired\. The reasoning traces may actually confuse the model and introduces error\. That said, the errors we observe suggest that a different kind ofLearn2Zinc\-CoTcould help\. Specifically,Learn2Zinc\-CoTfocused on mathematical and constraint\-level reasoning rather than simple step\-by\-step narration may reduce the semantic failures that current models make\. This is one possible future study\.

Regarding execution accuracy,error\-correction is clearly helpful for fine\-tuning\. TheLearn2Zinc\-Augmentedstrategy directly targets the main bottleneck of syntax errors\. By training on realistic error\-correction examples collected through cross\-model bootstrapping, S:< learn toboth generate code and fix common mistakes\. Using errors from multiple model sizes gives broad coverage of failure patterns\. It would be interesting to curate such fine\-tuning datasets for domain\-specific languages\.

Regarding solution accuracy, our ensemble that has 98% execution accuracy only achieves 35% solution accuracy\. This gap shows a clear split:while syntax is learnable, but constraint modeling remains hard for SLMs\. The models produce validMiniZinccode that compiles and runs, but often get the problem semantics wrong\. In our error anaylsis find two main categories of semantic error:

Type I Error: Contradictory constraints causing infeasibility\.Models introduce auxiliary variables with two incompatible definitions\. For example, a variable may be constrained to equal both the leftover supply and the truck count computed from the same shipped quantity\. The solver returns unsat even though the underlying problem is perfectly feasible\. An example of this type of error is given in Appendix[C\.1](https://arxiv.org/html/2607.20456#A3.SS1)\. As shown in Appendix\-Table[11](https://arxiv.org/html/2607.20456#A3.T11)\(Type I\), fixing these contradictions alone could improve solution accuracy by up to 18 percentage points depending on the model\.

Type II Error: Phantom variables forcing zero objectives\.Models create unnecessary decision variables \(for example, binary indicators in a pure LP\) and constrain them to equal expressions that exceed their declared bounds\. The only way to satisfy these constraints is to set all decision variables to zero, which gives a trivially wrong objective\. An example of this type of error is given in Appendix[C\.2](https://arxiv.org/html/2607.20456#A3.SS2)\. Appendix\-Table[11](https://arxiv.org/html/2607.20456#A3.T11)\(Type II\) shows this affects up to 20 percentage points\.

## 9Related Work

There is growing interest in leveraging LLMs for optimization tasks to transform the way that decision makers interact with solvers in the form of a co\-pilot system\(Simchi\-Leviet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib26), Wasserkruget al\.[2025](https://arxiv.org/html/2607.20456#bib.bib27), Kadıoğluet al\.[2026](https://arxiv.org/html/2607.20456#bib.bib42), Tsouroset al\.[2023](https://arxiv.org/html/2607.20456#bib.bib28)\)\.

#### From Natural Language to Optimization Models:

Earlier systems relied on rule\-based parsing to construct mixed\-integer or logic models \(e\.g\.,LGPSolverfor logic\-grid puzzles\(Jabrayilzade and Tekir[2020](https://arxiv.org/html/2607.20456#bib.bib32)\)andAutoLPfor linear programs\(Islamet al\.[2021](https://arxiv.org/html/2607.20456#bib.bib33)\)\)\. TheNl4Optcompetition\(Ramamonjisonet al\.[2022](https://arxiv.org/html/2607.20456#bib.bib16)\)formalized the task as a two\-step pipeline: named\-entity recognition followed by code generation\. Follow\-up works such asLaTeX2Solver\(Ramamonjisonet al\.[2023](https://arxiv.org/html/2607.20456#bib.bib34)\)extended parsing to mathematical documents, while the “Holy Grail 2\.0” blueprint\(Tsouroset al\.[2023](https://arxiv.org/html/2607.20456#bib.bib28)\)envisioned conversational assistants that refine models interactively\.

#### Prompt\-Driven and Learning\-based Copilots:

TheNer4Optline of works\(Dakleet al\.[2023](https://arxiv.org/html/2607.20456#bib.bib25), Kadıoğluet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib24)\)showed that fine\-tuning transformers on optimization\-specific corpora and inline entity tags boosts accuracy\. Retrieval\-augmented prompting specific to Constraint Programming settings has also proven effective\(Michailidiset al\.[2024](https://arxiv.org/html/2607.20456#bib.bib31)\)\. Modular agent pipelines push this further:ComplexORemploys a chain\-of\-experts architecture for difficult OR problems\(Xiaoet al\.[2023](https://arxiv.org/html/2607.20456#bib.bib17)\), whileOptiMUSdecomposes formulation, debugging, and solving into separate GPT agents\(AhmadiTeshniziet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib19)\)\. The recentGalaframework builds global agents for CP\(Caiet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib29)\)\. Fine\-tuning approaches includeOrlm\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\)andLllmOpt\(Jianget al\.[2025](https://arxiv.org/html/2607.20456#bib.bib30)\)\.

#### Code Generation for Rare Languages:

EsoLang\-Bench\(Sharma and Chopra[2026](https://arxiv.org/html/2607.20456#bib.bib39)\)shows frontier models drop from 85\-95% to 0\-11% on rare languages, with few\-shot and self\-reflection providing negligible benefit when pretraining coverage is absent\. Our work differs in that we study whether fine\-tuning, rather than prompting, can close this gap for constraint programming\.

#### Datasets and Evaluation:

Optimization\-centric corpora includeNl4Opt\(Ramamonjisonet al\.[2022](https://arxiv.org/html/2607.20456#bib.bib16)\),Nlp4Lp\(AhmadiTeshniziet al\.[2024](https://arxiv.org/html/2607.20456#bib.bib19)\),IndustryOr\(Huanget al\.[2024](https://arxiv.org/html/2607.20456#bib.bib23)\), andMILPsynthesis datasets\(Liet al\.[2023](https://arxiv.org/html/2607.20456#bib.bib35)\)\. Recent datasets includePlanetarium\(Zuoet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib36)\)for planning domain definition language \(PDDL\),Ehop\(Duchnowskiet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib37)\)for everyday NP\-Hard problems, andDualSchool\(Klamkinet al\.[2025](https://arxiv.org/html/2607.20456#bib.bib38)\)for optimization education\. Evaluation metrics have evolved from exact string match to execution and solution accuracy\.

Our work onLearn2Zincbuilds on theText2ZincdatasetSingirikondaet al\.\([2025](https://arxiv.org/html/2607.20456#bib.bib41)\)andText2ModelcopilotsKadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\), which together provide the data and baselines for the fine\-tuning study presented here\. The closest to our approach onLearn2Zincis the work onOptiMindZhanget al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib43)\)which fine\-tunes GPT\-OSS\-20B on optimization modeling for Gurobi modeling in Python\. Interestingly,Kadıoğluet al\.\([2026](https://arxiv.org/html/2607.20456#bib.bib42)\)reports that theMiniZinccapabilities ofOptiMindfine\-tuned GPT\-OSS\-20B is reduced to zero\. This is a call for future research on fine\-tuning models with potentially unintended consequences on losing capabilities in other domains\.

## 10Conclusion

We presentedLearn2Zinc,the first systematic study of fine\-tuning small language models \(SLMs\) for generatingMiniZincconstraint modelsfrom natural language\. Our results show thatMiniZincis genuinely out\-of\-distribution for current models, with near\-zero execution accuracy without fine\-tuning\. This highlights the limitations of prompt\-based approaches for domain\-specific languages with limited pretraining coverage\.

To address this, we introducedseveral fine\-tuning strategiesincluding cross\-model error bootstrapping, a method for constructing realistic syntax error\-correction datasets by leveraging failures across multiple models\. Combined with grammar\-based synthetic corruptions, this approach enables models to learn both code generation and targeted error correction\. We further show that, unlike in prompting settings, chain\-of\-thought fine\-tuning does not improve performance when foundational syntax knowledge is absent\.

OurLearn2Zinc\-Augmentedstrategy significantly improves execution accuracy across all model sizes of SLMs\. When combined with self\-reflection and a simple top\-down ensemble,we achieve 98% execution accuracy, effectively resolving the syntax bottleneck forMiniZinc\. However, solution accuracy saturates at 34%, revealing a clear gap between syntactic correctness and semantic reasoning\. Our initial error analysis attributes this gap to recurring issues such as contradictory constraints and phantom variables\. These findings suggest that while syntax can be learned efficiently through targeted fine\-tuning, constraint reasoning remains the primary challenge for SLM\-based text\-to\-model translation\.

Future work should focus onclosing the execution\-solution gapthrough improved supervision and reasoning support\. Promising directions include training with structured reasoning traces aligned with constraint modeling subtasks, leveraging intermediate supervision forMiniZincgeneration, and distilling reasoning capabilities from larger models\. In addition, integrating fine\-tuned models with grammar\-constrained decoding, prompting strategies, and modular or agentic pipelines may further improve solution quality\.

## References

- OptiMUS\-0\.3: using large language models to model and solve optimization problems at scale\.External Links:2407\.19633,[Link](https://arxiv.org/abs/2407.19633)Cited by:[§2\.2](https://arxiv.org/html/2607.20456#S2.SS2.p1.5),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- M\. R\. Bussieck and A\. Meeraus \(2004\)General algebraic modeling system \(gams\)\.InModeling Languages in Mathematical Optimization,pp\. 137–157\.External Links:[Document](https://dx.doi.org/10.1007/978-1-4613-0215-5%5F8),ISBN 978\-1\-4613\-0215\-5Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p2.1)\.
- J\. Cai, S\. Kadıoğlu, and B\. Dilkina \(2025\)Gala: global llm agents for text\-to\-model translation\.External Links:2509\.08970,[Link](https://arxiv.org/abs/2509.08970)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- P\. P\. Dakle, S\. Kadıoğlu, K\. Uppuluri, R\. Politi, P\. Raghavan, S\. Rallabandi, and R\. Srinivasamurthy \(2023\)Ner4opt: named entity recognition for optimization modelling from natural language\.InInternational Conference on Integration of Constraint Programming, Artificial Intelligence, and Operations Research,pp\. 299–319\.Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- DeepSeek\-AIet al\.\(2025\)DeepSeek\-v3 technical report\.External Links:2412\.19437,[Link](https://arxiv.org/abs/2412.19437)Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1)\.
- A\. Duchnowski, E\. Pavlick, and A\. Koller \(2025\)EHOP: a dataset of everyday np\-hard optimization problems\.External Links:2502\.13776,[Link](https://arxiv.org/abs/2502.13776)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- T\. Guns \(2019\)Increasing modeling language convenience with a universal n\-dimensional array, cppy as python\-embedded example\.InProceedings of the 18th workshop on Constraint Modelling and Reformulation at CP \(Modref 2019\),Vol\.19\.Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p2.1)\.
- C\. Huang, Z\. Tang, D\. Ge, S\. Hu, R\. Jiang, B\. Wang, Z\. Wang, and X\. Zheng \(2024\)ORLM: a customizable framework in training large models for automated optimization modeling\.External Links:2405\.17743,[Link](https://arxiv.org/abs/2405.17743)Cited by:[Table 1](https://arxiv.org/html/2607.20456#S3.T1.1.2.1.1),[§3](https://arxiv.org/html/2607.20456#S3.p2.1),[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p1.1),[§7](https://arxiv.org/html/2607.20456#S7.p1.1),[§7](https://arxiv.org/html/2607.20456#S7.p2.1),[§7](https://arxiv.org/html/2607.20456#S7.p6.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- Md\. S\. Islam, F\. Mamud, R\. U\. Haque, and A\. Y\. Saber \(2021\)Automatic formulation and optimization of linear problems from a structured paragraph\.InProc\. 2021 Int\. Conf\. on Science & Contemporary Technologies \(ICSCT\),External Links:[Document](https://dx.doi.org/10.1109/ICSCT53883.2021.9642516)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px1.p1.1)\.
- E\. Jabrayilzade and S\. Tekir \(2020\)LGPSolver – solving logic grid puzzles automatically\.InFindings of the Association for Computational Linguistics: EMNLP 2020,pp\. 1118–1123\.External Links:[Document](https://dx.doi.org/10.18653/v1/2020.findings-emnlp.100)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px1.p1.1)\.
- C\. Jiang, X\. Shu, H\. Qian, X\. Lu, J\. Zhou, A\. Zhou, and Y\. Yu \(2025\)LLMOPT: learning to define and solve general optimization problems from scratch\.InProceedings of the 13th International Conference on Learning Representations \(ICLR\),Singapore\.External Links:[Link](https://openreview.net/forum?id=9OMvtboTJg)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- S\. Kadıoğlu, P\. Pravin Dakle, K\. Uppuluri, R\. Politi, P\. Raghavan, S\. Rallabandi, and R\. Srinivasamurthy \(2024\)Ner4Opt: named entity recognition for optimization modelling from natural language\.Constraints,pp\. 1–39\.Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1),[§2\.2](https://arxiv.org/html/2607.20456#S2.SS2.p1.5),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- S\. Kadıoğlu, K\. Uppuluri, and A\. Singirikonda \(2026\)Modeling copilots for text\-to\-model translation\.External Links:2604\.12955,[Link](https://arxiv.org/abs/2604.12955)Cited by:[§2\.3](https://arxiv.org/html/2607.20456#S2.SS3.p1.1),[§4\.3](https://arxiv.org/html/2607.20456#S4.SS3.p1.1),[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p1.1),[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p3.1),[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p4.1),[Table 2](https://arxiv.org/html/2607.20456#S4.T2),[Table 2](https://arxiv.org/html/2607.20456#S4.T2.1.27.2),[§6\.2](https://arxiv.org/html/2607.20456#S6.SS2.p2.1),[Table 8](https://arxiv.org/html/2607.20456#S6.T8.1.7.2),[§7](https://arxiv.org/html/2607.20456#S7.p5.1),[§7](https://arxiv.org/html/2607.20456#S7.p6.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p2.1),[§9](https://arxiv.org/html/2607.20456#S9.p1.1)\.
- M\. Klamkin, A\. Deza, S\. Cheng, H\. Zhao, and P\. V\. Hentenryck \(2025\)DualSchool: how reliable are llms for optimization education?\.External Links:2505\.21775,[Link](https://arxiv.org/abs/2505.21775)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- Q\. Li, L\. Zhang, and V\. Mak\-Hau \(2023\)Synthesizing mixed\-integer linear programming models from natural language descriptions\.External Links:2311\.15271,[Link](https://arxiv.org/abs/2311.15271)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- K\. Liang, Y\. Lu, J\. Mao, S\. Sun, C\. Yang, C\. Zeng, X\. Jin, H\. Qin, R\. Zhu, and C\. Teo \(2026\)Large\-scale optimization model auto\-formulation: harnessing llm flexibility via structured workflow\.External Links:2601\.09635,[Link](https://arxiv.org/abs/2601.09635)Cited by:[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p1.1),[§6\.2](https://arxiv.org/html/2607.20456#S6.SS2.p4.1),[Table 9](https://arxiv.org/html/2607.20456#S7.T9),[§7](https://arxiv.org/html/2607.20456#S7.p1.1),[§7](https://arxiv.org/html/2607.20456#S7.p2.1),[§7](https://arxiv.org/html/2607.20456#S7.p5.1),[§7](https://arxiv.org/html/2607.20456#S7.p6.1)\.
- K\. Michailidis, D\. Tsouros, and T\. Guns \(2024\)Constraint Modelling with LLMs Using In\-Context Learning\.In30th International Conference on Principles and Practice of Constraint Programming \(CP 2024\),Leibniz International Proceedings in Informatics \(LIPIcs\), Vol\.307,pp\. 20:1–20:27\.External Links:[Document](https://dx.doi.org/10.4230/LIPIcs.CP.2024.20)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- N\. Nethercote, P\. J\. Stuckey, R\. Becket, S\. Brand, G\. J\. Duck, and G\. Tack \(2007\)MiniZinc: towards a standard cp modelling language\.InPrinciples and Practice of Constraint Programming – CP 2007,C\. Bessière \(Ed\.\),Berlin, Heidelberg,pp\. 529–543\.External Links:ISBN 978\-3\-540\-74970\-7Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p2.1),[§2\.1](https://arxiv.org/html/2607.20456#S2.SS1.p1.1)\.
- OpenAIet al\.\(2024\)GPT\-4 technical report\.External Links:2303\.08774,[Link](https://arxiv.org/abs/2303.08774)Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1)\.
- R\. Ramamonjison, T\. Yu, L\. Xing, M\. Mostajabdaveh, X\. Li, X\. Fu, X\. Han, Y\. Chen, R\. Li, K\. Mao, and Y\. Zhang \(2023\)LaTeX2Solver: a hierarchical semantic parsing of LaTeX document into code for an assistive optimization modeling application\.InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics \(Volume 3: System Demonstrations\),pp\. 471–478\.External Links:[Document](https://dx.doi.org/10.18653/v1/2023.acl-demo.45)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px1.p1.1)\.
- R\. Ramamonjison, H\. Li, T\. T\. Yu, S\. He, V\. Rengan, A\. Banitalebi\-Dehkordi, Z\. Zhou, and Y\. Zhang \(2022\)Augmenting operations research with auto\-formulation of optimization models from problem descriptions\.arXiv\.External Links:[Document](https://dx.doi.org/10.48550/ARXIV.2209.15565),[Link](https://arxiv.org/abs/2209.15565)Cited by:[§2\.2](https://arxiv.org/html/2607.20456#S2.SS2.p1.5),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px1.p1.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.
- A\. Sharma and P\. Chopra \(2026\)EsoLang\-bench: evaluating genuine reasoning in large language models via esoteric programming languages\.External Links:2603\.09678,[Link](https://arxiv.org/abs/2603.09678)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px3.p1.1)\.
- D\. Simchi\-Levi, T\. Dai, I\. Menache, and M\. X\. Wu \(2025\)Democratizing optimization with generative ai\.Johns Hopkins Carey Business School Research Paper Forthcoming\.Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1),[§9](https://arxiv.org/html/2607.20456#S9.p1.1)\.
- A\. Singirikonda, S\. Kadıoğlu, and K\. Uppuluri \(2025\)Text2Zinc: a cross\-domain dataset for modeling optimization and satisfaction problems in minizinc\.External Links:2503\.10642,[Link](https://arxiv.org/abs/2503.10642)Cited by:[§2\.2](https://arxiv.org/html/2607.20456#S2.SS2.p1.5),[Table 1](https://arxiv.org/html/2607.20456#S3.T1.1.3.1.1),[§4\.4](https://arxiv.org/html/2607.20456#S4.SS4.p4.1),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p2.1)\.
- G\. Teamet al\.\(2025\)Gemini: a family of highly capable multimodal models\.External Links:2312\.11805,[Link](https://arxiv.org/abs/2312.11805)Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1)\.
- D\. Tsouros, H\. Verhaeghe, S\. Kadıoğlu, and T\. Guns \(2023\)Holy grail 2\.0: from natural language to constraint models\.arXiv preprint arXiv:2308\.01589\.Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px1.p1.1),[§9](https://arxiv.org/html/2607.20456#S9.p1.1)\.
- S\. Wasserkrug, L\. Boussioux, D\. den Hertog, F\. Mirzazadeh, S\. I\. Birbil, J\. Kurtz, and D\. Maragno \(2025\)Enhancing decision making through the integration of large language models and operations research optimization\.Proceedings of the AAAI Conference on Artificial Intelligence39\(27\),pp\. 28643–28650\.External Links:[Document](https://dx.doi.org/10.1609/aaai.v39i27.35090)Cited by:[§1](https://arxiv.org/html/2607.20456#S1.p3.1),[§9](https://arxiv.org/html/2607.20456#S9.p1.1)\.
- Z\. Xiao, D\. Zhang, Y\. Wu, L\. Xu, Y\. J\. Wang, X\. Han, X\. Fu, T\. Zhong, J\. Zeng, M\. Song,et al\.\(2023\)Chain\-of\-experts: when llms meet complex operations research problems\.InThe Twelfth International Conference on Learning Representations,Cited by:[§2\.2](https://arxiv.org/html/2607.20456#S2.SS2.p1.5),[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px2.p1.1)\.
- X\. Zhang, Z\. Chen, H\. Zope, H\. Barbalho, K\. Mellou, M\. Molinaro, J\. Kulkarni, I\. Menache, and S\. Li \(2026\)OptiMind: teaching llms to think like optimization experts\.External Links:2509\.22979,[Link](https://arxiv.org/abs/2509.22979)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p2.1.5)\.
- M\. Zuo, F\. P\. Velez, X\. Li, M\. L\. Littman, and S\. H\. Bach \(2025\)Planetarium: a rigorous benchmark for translating text to structured planning languages\.External Links:2407\.03321,[Link](https://arxiv.org/abs/2407.03321)Cited by:[§9](https://arxiv.org/html/2607.20456#S9.SS0.SSS0.Px4.p1.1)\.

## Appendix AFine\-tuning Examples

### Learn2Zinc\-BaseExample

Instruction:GenerateMiniZinccode for the following optimization problem\.

A rescue material distribution company needs to allocate rescue materials to three disaster\-stricken areas: A, B, and C\. These places have different degrees of disaster, so the required amount of rescue materials is also different\. The goal is to minimize the total distribution time associated with these locations, with a distribution time weight of 50 hours for A, 100 hours for B, and 200 hours for C\.

The allocation must adhere to the following conditions:

\- The total amount of rescue materials for locations A and B cannot exceed 500 units\. \- The total amount of rescue materials for locations B and C must be at least 300 units\. \- The amount of rescue materials for location C cannot exceed the maximum demand of 600 units\. \- The amount of rescue materials for location A must be at least 50 units\.

Response:

int:wA=50;

int:wB=100;

int:wC=200;

int:A\_min=50;

int:A\_max=500;

int:B\_min=0;

int:B\_max=400;

int:C\_min=0;

int:C\_max=600;

int:AB\_max=500;

int:BC\_min=300;

varA\_min\.\.A\_max:A;

varB\_min\.\.B\_max:B;

varC\_min\.\.C\_max:C;

var0\.\.\(wA\*A\_max\+wB\*B\_max\+wC\*C\_max\):total\_time;

constraintA\+B<=AB\_max;

constraintB\+C\>=BC\_min;

constrainttotal\_time=wA\*A\+wB\*B\+wC\*C;

solveminimizetotal\_time;

### Learn2Zinc\-CoTExample

Instruction:Generate MiniZinc code for the following optimization problem\.First outline your reasoning, then provide the code\.A rescue material distribution company needs to allocate rescue materials to three disaster\-stricken areas: A, B, and C\. These places have different degrees of disaster, so the required amount of rescue materials is also different\. The goal is to minimize the total distribution time associated with these locations, with a distribution time weight of 50 hours for A, 100 hours for B, and 200 hours for C\.

The allocation must adhere to the following conditions:

\- The total amount of rescue materials for locations A and B cannot exceed 500 units\. \- The total amount of rescue materials for locations B and C must be at least 300 units\. \- The amount of rescue materials for location C cannot exceed the maximum demand of 600 units\. \- The amount of rescue materials for location A must be at least 50 units\.

Considering that each location has a specific demand for rescue materials \(A: 0\-500, B: 0\-400, C: 0\-600\), and due to the limitation of the quantity of materials, since we cannot have fractional units of rescue materials, what is the minimum total distribution time in hours for the optimal allocation of rescue materials to the three locations within the specified demand ranges? Please round the answer to the nearest hour\.

Response:

Reasoning:Variables: A

∈\\in\[50, 500\], B

∈\\in\[0, 400\], C

∈\\in\[0, 600\]

Constraints: A \+ B

≤\\leq500, B \+ C

≥\\geq300

Objective: minimize total\_time = 50\*A \+ 100\*B \+ 200\*C

Code:

int:wA=50;

int:wB=100;

int:wC=200;

int:A\_min=50;

int:A\_max=500;

int:B\_min=0;

int:B\_max=400;

int:C\_min=0;

int:C\_max=600;

int:AB\_max=500;

int:BC\_min=300;

varA\_min\.\.A\_max:A;

varB\_min\.\.B\_max:B;

varC\_min\.\.C\_max:C;

var0\.\.\(wA\*A\_max\+wB\*B\_max\+wC\*C\_max\):total\_time;

constraintA\+B<=AB\_max;

constraintB\+C\>=BC\_min;

constrainttotal\_time=wA\*A\+wB\*B\+wC\*C;

solveminimizetotal\_time;

### Syntax Correction Example

Instruction:Fix the following MiniZinc code that has syntax errors\.

Problem: A shipping company needs to transport goods from a source port to a destination port\. There are 3 ships available for selection: Ship A, B, and C\. The cargo capacity, transportation cost, and transportation time for each ship are shown in the table below:

ShipCapacity \(tons\)Cost \(M yuan\)Time \(days\)A10001010B80088C1200127Now, there is a need to transport 3000 tons of cargo and it must arrive within 15 days\. How should the ships be scheduled for transportation to minimize the total transportation cost? Each ship can only make one trip, and the amount of cargo transported cannot exceed its cargo capacity\.

Broken code:

int:NSHIPS=3;

setofint:SHIPS=1\.\.NSHIPS;

array\[SHIPS\]ofstring:ship\_name=\["A","B","C"\];

array\[SHIPS\]ofint:cap=\[1000,800,1200\];

array\[SHIPS\]ofint:cost=\[10,8,12\];

array\[SHIPS\]ofint:time=\[10,8,7\];

int:demand=3000;

int:max\_days=15;

int:max\_trips=ceil\(demanddivmin\(cap\)\);

array\[SHIPS\]ofvar0\.\.max\_trips:trips;

var0\.\.demand:cargo\_each;

array\[SHIPS\]ofvar0\.\.demand:x;

var0\.\.\(max\_trips\*max\(cost\)\):total\_cost;

constraintforall\(iinSHIPS\)\(

cargo\_each\[i\]<=cap\[i\]/\\

x\[i\]=trips\[i\]\*cargo\_each\[i\]/\\

x\[i\]<=cap\[i\]/\\

x\[i\]\>=0

\);

constraintsum\(iinSHIPS\)\(x\[i\]\)=demand;

constraintforall\(iinSHIPS\)\(

cargo\_each\[i\]<=cap\[i\]/\\

x\[i\]<=cap\[i\]/\\

x\[i\]\>=0

\);

constraintforall\(iinSHIPS\)\(

time\[i\]<=max\_days/\\

time\[i\]\>=0

\);

constrainttotal\_cost=sum\(iinSHIPS\)\(cost\[i\]\*trips\[i\]\);

solveminimizetotal\_cost;

Error:type error: array access attempted on expression of type ‘var int’

Response:

int:N=3;

setofint:SHIPS=1\.\.N;

array\[SHIPS\]ofstring:name=\["A","B","C"\];

array\[SHIPS\]ofint:capacity=\[1000,800,1200\];

array\[SHIPS\]ofint:cost=\[10,8,12\];

array\[SHIPS\]ofint:time=\[10,8,7\];

int:demand=3000;

int:max\_days=15;

array\[SHIPS\]ofvar0\.\.1:make\_trip;

array\[SHIPS\]ofvar0\.\.demand:cargo;

var0\.\.\(max\_days\*max\(cost\)\):total\_cost;

constraintforall\(iinSHIPS\)\(

cargo\[i\]<=capacity\[i\]\*make\_trip\[i\]

\);

constraintsum\(iinSHIPS\)\(cargo\[i\]\)=demand;

constraintforall\(iinSHIPS\)\(

make\_trip\[i\]=bool2int\(time\[i\]\>=1\)

\);

constraintsum\(iinSHIPS\)\(time\[i\]\*make\_trip\[i\]\)<=max\_days;

constrainttotal\_cost=sum\(iinSHIPS\)\(cost\[i\]\*make\_trip\[i\]\);

solveminimizetotal\_cost;

## Appendix BTraining Details

Table 10:Training hyperparameters for each model\. All models used max sequence length of 4,096 tokens and batch size of 2\.ModelLoRArrLoRAα\\alphaLREpochsQuantizationQwen3\-0\.6B64642e\-438\-bitLLaMA\-3\.2\-1B64642e\-438\-bitLLaMA\-3\.2\-3B64642e\-438\-bitGemma\-2\-9B64641e\-438\-bitGPT\-OSS\-20B32642e\-434\-bit
## Appendix CError Analysis

### C\.1Type I: Contradictory Constraints

LLM\-Generated Constraint Errors in MiniZinc

\(a\) LLM\-Generated Code Result: UNSATISFIABLE[⬇](data:text/plain;base64,JSBQYXJhbWV0ZXJzIChlbWJlZCBhbGwgZGF0YSB2YWx1ZXMgZGlyZWN0bHkgaGVyZSkKaW50OiBXID0gNjsKaW50OiBQID0gNTsKCnNldCBvZiBpbnQ6IFdhcmVob3VzZXMgPSAxLi5XOwpzZXQgb2YgaW50OiBQb3J0cyA9IDEuLlA7CgphcnJheVtXYXJlaG91c2VzXSBvZiBzdHJpbmc6IHduYW1lID0KICBbIlZlcm9uYSIsIlBlcnVnaWEiLCJSb21lIiwKICAgIlBlc2NhcmEiLCJUYXJhbnRvIiwiTGFtZXppYSJdOwphcnJheVtQb3J0c10gb2Ygc3RyaW5nOiBwbmFtZSA9CiAgWyJHZW5vYSIsIlZlbmljZSIsIkFuY29uYSIsIk5hcGxlcyIsIkJhcmkiXTsKCmFycmF5W1dhcmVob3VzZXNdIG9mIGludDogc3VwcGx5ID0KICBbMTAsIDEyLCAyMCwgMjQsIDE4LCA0MF07CmFycmF5W1BvcnRzXSBvZiBpbnQ6IGRlbWFuZCA9CiAgWzIwLCAxNSwgMjUsIDMzLCAyMV07CgolIFRydWNrIGNhcGFjaXR5OiAyIGNvbnRhaW5lcnMgcGVyIHRydWNrCmludDogY2FwID0gMjsKCiUgRGlzdGFuY2UgKGttKSBwZXIgY29udGFpbmVyIChmb3IgY29zdCBjYWxjdWxhdGlvbikKYXJyYXlbV2FyZWhvdXNlcywgUG9ydHNdIG9mIGludDogZGlzdCA9CiAgW3wgMjkwLCAxMTUsIDM1NSwgNzE1LCA4MTAKICAgfCAzODAsIDM0MCwgMTY1LCAzODAsIDYxMAogICB8IDUwNSwgNTMwLCAyODUsIDIyMCwgNDUwCiAgIHwgNjU1LCA0NTAsIDE1NSwgMjQwLCAzMTUKICAgfCAxMDEwLCA4NDAsIDU1MCwgMzA1LCA5NQogICB8IDEwNzIsIDEwOTcsIDc0NywgMzcyLCAzMzMgfF07CgppbnQ6IGNvc3RfcGVyX2ttID0gMzA7)int:W=6;int:P=5;setofint:Warehouses=1\.\.W;setofint:Ports=1\.\.P;array\[Warehouses\]ofstring:wname=\["Verona","Perugia","Rome","Pescara","Taranto","Lamezia"\];array\[Ports\]ofstring:pname=\["Genoa","Venice","Ancona","Naples","Bari"\];array\[Warehouses\]ofint:supply=\[10,12,20,24,18,40\];array\[Ports\]ofint:demand=\[20,15,25,33,21\];int:cap=2;array\[Warehouses,Ports\]ofint:dist=\[\|290,115,355,715,810\|380,340,165,380,610\|505,530,285,220,450\|655,450,155,240,315\|1010,840,550,305,95\|1072,1097,747,372,333\|\];int:cost\_per\_km=30;[⬇](data:text/plain;base64,JSBWYXJpYWJsZXMKYXJyYXlbV2FyZWhvdXNlcywgUG9ydHNdIG9mIHZhciAwLi40MDogeDsKYXJyYXlbV2FyZWhvdXNlc10gb2YgdmFyIDAuLjQwOiB1c2VkOwphcnJheVtXYXJlaG91c2VzXSBvZiB2YXIgMC4uNDA6IHNoaXBwZWQ7CnZhciAwLi4xMDAwMDAwOiB0b3RhbF9jb3N0Ow==)array\[Warehouses,Ports\]ofvar0\.\.40:x;array\[Warehouses\]ofvar0\.\.40:used;array\[Warehouses\]ofvar0\.\.40:shipped;var0\.\.1000000:total\_cost;[⬇](data:text/plain;base64,JSBDb25zdHJhaW50cwpjb25zdHJhaW50CiAgZm9yYWxsKGkgaW4gV2FyZWhvdXNlcykgKAogICAgc2hpcHBlZFtpXSA9IHN1bShqIGluIFBvcnRzKSh4W2ksal0pCiAgKTs=)constraintforall\(iinWarehouses\)\(shipped\[i\]=sum\(jinPorts\)\(x\[i,j\]\)\);[⬇](data:text/plain;base64,Y29uc3RyYWludCAgICAgICAgICAgICAgICAgICAgICAgICAgJSBCVUcgMQogIGZvcmFsbChpIGluIFdhcmVob3VzZXMpICgKICAgIHNoaXBwZWRbaV0gPSBzdXBwbHlbaV0gLSB1c2VkW2ldCiAgKTs=)constraintforall\(iinWarehouses\)\(shipped\[i\]=supply\[i\]\-used\[i\]\);[⬇](data:text/plain;base64,Y29uc3RyYWludAogIGZvcmFsbChqIGluIFBvcnRzKSAoCiAgICBzdW0oaSBpbiBXYXJlaG91c2VzKSh4W2ksal0pID0gZGVtYW5kW2pdCiAgKTs=)constraintforall\(jinPorts\)\(sum\(iinWarehouses\)\(x\[i,j\]\)=demand\[j\]\);[⬇](data:text/plain;base64,JSBFYWNoIHRydWNrIGNhbiBjYXJyeSB1cCB0byAyICAgICAgJSBCVUcgMgolIGNvbnRhaW5lcnM7IGNvdW50IHRydWNrcwpjb25zdHJhaW50CiAgZm9yYWxsKGkgaW4gV2FyZWhvdXNlcykgKAogICAgdXNlZFtpXSA9IHN1bShqIGluIFBvcnRzKQogICAgICAgICAgICAgICh4W2ksal0pIGRpdiBjYXAKICApOw==)constraintforall\(iinWarehouses\)\(used\[i\]=sum\(jinPorts\)\(x\[i,j\]\)divcap\);[⬇](data:text/plain;base64,JSBUb3RhbCBjb3N0ICAgICAgICAgICAgICAgICAgICAgICAgJSBCVUcgMwpjb25zdHJhaW50CiAgdG90YWxfY29zdCA9IGNvc3RfcGVyX2ttCiAgICAqIHN1bShpIGluIFdhcmVob3VzZXMpKHVzZWRbaV0pCiAgICAqIGNhcDs=)constrainttotal\_cost=cost\_per\_km\*sum\(iinWarehouses\)\(used\[i\]\)\*cap;[⬇](data:text/plain;base64,JSBPYmplY3RpdmUKc29sdmUgbWluaW1pemUgdG90YWxfY29zdDs=)solveminimizetotal\_cost;

\(b\) Corrected Code Optimal cost: 904590[⬇](data:text/plain;base64,JSBQYXJhbWV0ZXJzIChlbWJlZCBhbGwgZGF0YSB2YWx1ZXMgZGlyZWN0bHkgaGVyZSkKaW50OiBXID0gNjsKaW50OiBQID0gNTsKc2V0IG9mIGludDogV2FyZWhvdXNlcyA9IDEuLlc7CnNldCBvZiBpbnQ6IFBvcnRzID0gMS4uUDsKYXJyYXlbV2FyZWhvdXNlc10gb2Ygc3RyaW5nOiB3bmFtZSA9CiAgWyJWZXJvbmEiLCJQZXJ1Z2lhIiwiUm9tZSIsCiAgICJQZXNjYXJhIiwiVGFyYW50byIsIkxhbWV6aWEiXTsKYXJyYXlbUG9ydHNdIG9mIHN0cmluZzogcG5hbWUgPQogIFsiR2Vub2EiLCJWZW5pY2UiLCJBbmNvbmEiLCJOYXBsZXMiLCJCYXJpIl07CmFycmF5W1dhcmVob3VzZXNdIG9mIGludDogc3VwcGx5ID0KICBbMTAsIDEyLCAyMCwgMjQsIDE4LCA0MF07CmFycmF5W1BvcnRzXSBvZiBpbnQ6IGRlbWFuZCA9CiAgWzIwLCAxNSwgMjUsIDMzLCAyMV07CiUgVHJ1Y2sgY2FwYWNpdHk6IDIgY29udGFpbmVycyBwZXIgdHJ1Y2sKaW50OiBjYXAgPSAyOwolIERpc3RhbmNlIChrbSkKYXJyYXlbV2FyZWhvdXNlcywgUG9ydHNdIG9mIGludDogZGlzdCA9CiAgW3wgMjkwLCAxMTUsIDM1NSwgNzE1LCA4MTAKICAgfCAzODAsIDM0MCwgMTY1LCAzODAsIDYxMAogICB8IDUwNSwgNTMwLCAyODUsIDIyMCwgNDUwCiAgIHwgNjU1LCA0NTAsIDE1NSwgMjQwLCAzMTUKICAgfCAxMDEwLCA4NDAsIDU1MCwgMzA1LCA5NQogICB8IDEwNzIsIDEwOTcsIDc0NywgMzcyLCAzMzMgfF07)int:W=6;int:P=5;setofint:Warehouses=1\.\.W;setofint:Ports=1\.\.P;array\[Warehouses\]ofstring:wname=\["Verona","Perugia","Rome","Pescara","Taranto","Lamezia"\];array\[Ports\]ofstring:pname=\["Genoa","Venice","Ancona","Naples","Bari"\];array\[Warehouses\]ofint:supply=\[10,12,20,24,18,40\];array\[Ports\]ofint:demand=\[20,15,25,33,21\];int:cap=2;array\[Warehouses,Ports\]ofint:dist=\[\|290,115,355,715,810\|380,340,165,380,610\|505,530,285,220,450\|655,450,155,240,315\|1010,840,550,305,95\|1072,1097,747,372,333\|\];[⬇](data:text/plain;base64,JSBDb3N0ID0gMzAgZXVyb3MgcGVyIGttIHBlciBDT05UQUlORVIKaW50OiBjb3N0X3Blcl9rbSA9IDMwOwolIFRoZSBjb3N0IHRvIHRyYW5zcG9ydCBlYWNoIGNvbnRhaW5lciBvbgolIHJvdXRlIChpLGopIGlzIGNvc3RfcGVyX2ttICogZGlzdFtpLGpdCmFycmF5W1dhcmVob3VzZXMsIFBvcnRzXSBvZiBpbnQ6IGNvc3QgPQogIGFycmF5MmQoV2FyZWhvdXNlcywgUG9ydHMsCiAgW2Nvc3RfcGVyX2ttICogZGlzdFtpLGpdCiAgIHwgaSBpbiBXYXJlaG91c2VzLCBqIGluIFBvcnRzXSk7CiUgVmFyaWFibGVzCmFycmF5W1dhcmVob3VzZXMsIFBvcnRzXSBvZiB2YXIgMC4uNDA6IHg7CnZhciAwLi4xMDAwMDAwMDogdG90YWxfY29zdDs=)int:cost\_per\_km=30;array\[Warehouses,Ports\]ofint:cost=array2d\(Warehouses,Ports,\[cost\_per\_km\*dist\[i,j\]\|iinWarehouses,jinPorts\]\);array\[Warehouses,Ports\]ofvar0\.\.40:x;var0\.\.10000000:total\_cost;[⬇](data:text/plain;base64,JSBDb25zdHJhaW50cwolIFN1cHBseSBjb25zdHJhaW50OiBjYW5ub3Qgc2hpcCBtb3JlCiUgdGhhbiBhdmFpbGFibGUKY29uc3RyYWludCAgICAgICAgICAgICAgICAgICAgICAgICAgJSBGSVggMQogIGZvcmFsbChpIGluIFdhcmVob3VzZXMpICgKICAgIHN1bShqIGluIFBvcnRzKSh4W2ksal0pIDw9IHN1cHBseVtpXQogICk7)constraintforall\(iinWarehouses\)\(sum\(jinPorts\)\(x\[i,j\]\)<=supply\[i\]\);[⬇](data:text/plain;base64,JSBEZW1hbmQgY29uc3RyYWludDogZWFjaCBwb3J0IG11c3QKJSByZWNlaXZlIGV4YWN0bHkgaXRzIGRlbWFuZApjb25zdHJhaW50CiAgZm9yYWxsKGogaW4gUG9ydHMpICgKICAgIHN1bShpIGluIFdhcmVob3VzZXMpKHhbaSxqXSkgPSBkZW1hbmRbal0KICApOw==)constraintforall\(jinPorts\)\(sum\(iinWarehouses\)\(x\[i,j\]\)=demand\[j\]\);[⬇](data:text/plain;base64,JSBUb3RhbCBjb3N0OiAzMCBldXJvcyAqIGRpc3RhbmNlICAgJSBGSVggMiwzCiUgKiBudW1iZXIgb2YgY29udGFpbmVycyBvbiBlYWNoIHJvdXRlCmNvbnN0cmFpbnQKICB0b3RhbF9jb3N0ID0gc3VtKGkgaW4gV2FyZWhvdXNlcywKICAgICAgICAgICAgICAgICAgIGogaW4gUG9ydHMpCiAgICAgICAgICAgICAgIChjb3N0W2ksal0gKiB4W2ksal0pOw==)constrainttotal\_cost=sum\(iinWarehouses,jinPorts\)\(cost\[i,j\]\*x\[i,j\]\);[⬇](data:text/plain;base64,JSBPYmplY3RpdmUKc29sdmUgbWluaW1pemUgdG90YWxfY29zdDs=)solveminimizetotal\_cost;

Key ErrorsContradictory constraints \(causes UNSATISFIABLE\)\.Bugs 1 and 2 jointly forceused\[i\]to satisfy two incompatible definitions:used\[i\] = supply\[i\]−\-shipped\[i\]andused\[i\] = shipped\[i\] div 2\. Substituting, this requiressupply\[i\]−\-shipped\[i\] = shipped\[i\] div 2for every warehouse, which has no valid integer solution for most input values\. The solver correctly reports the model as infeasible\. The fix removesused\[i\]andshipped\[i\]entirely and replaces them with a single inequalitysum\(x\[i,j\]\) <= supply\[i\]\.Wrong objective formula\.Bug 3 computes cost ascost\_per\_km \* sum\(used\[i\]\) \* cap, which \(a\) never references thedistmatrix, treating all routes as equally expensive, and \(b\) models cost per truck rather than per container\. The problem states the rate is 30 euros/km*per container*, so the correct objective is simplysum\(cost\_per\_km \* dist\[i,j\] \* x\[i,j\]\)over all pairs, with no dependence on truck capacity\.

### C\.2Type II: Phantom Variables

LLM\-Generated Phantom Variable and Index Errors in MiniZinc

\(a\) LLM\-Generated Code — Result: obj = 0[⬇](data:text/plain;base64,JSBQYXJhbWV0ZXJzIChlbWJlZCBhbGwgZGF0YSB2YWx1ZXMgZGlyZWN0bHkgaGVyZSkKaW50OiBOID0gMzsKc2V0IG9mIGludDogUCA9IDEuLk47CmludDogTSA9IDM7CnNldCBvZiBpbnQ6IEUgPSAxLi5NOwoKJSBFcXVpcG1lbnQgY29kZXM6IDE9QSwgMj1CLCAzPUMKYXJyYXlbRV0gb2YgaW50OiBhdmFpbCA9IFszMDAsIDQwMCwgNDIwXTsKCiUgUHJvZml0IHBlciB1bml0IChwZXIgdGhvdXNhbmQgeXVhbik6CiUgWzMsIDIsIDIuOV0KYXJyYXlbUF0gb2YgZmxvYXQ6IHByb2ZpdCA9CiAgWzMuMCwgMi4wLCAyLjldOw==)int:N=3;setofint:P=1\.\.N;int:M=3;setofint:E=1\.\.M;array\[E\]ofint:avail=\[300,400,420\];array\[P\]offloat:profit=\[3\.0,2\.0,2\.9\];[⬇](data:text/plain;base64,JSBVc2FnZSBjb2VmZmljaWVudHMgcGVyIHVuaXQ6ICAgICAgICUgQlVHIDEKJSBbOCwgMiwgMTBdIGZvciBJLCBJSSwgSUlJIG9uIEEsQixDCmFycmF5W1AsIEVdIG9mIGludDogdXNlID0KICBhcnJheTJkKFAsIEUsClsKICA4LCAyLCAxMCwgICAlIEkgb24gQSxCLEMKICAxMCwgNSwgOCwgICAlIElJIG9uIEEsQixDCiAgMiwgMTMsIDEwICAgJSBJSUkgb24gQSxCLEMKXSk7)array\[P,E\]ofint:use=array2d\(P,E,\[8,2,10,10,5,8,2,13,10\]\);[⬇](data:text/plain;base64,JSBWYXJpYWJsZXMKJSB4W3BdID0gcHJvZHVjdGlvbiBxdWFudGl0eSBvZiBwcm9kdWN0IHAKJSAoaW4gInRob3VzYW5kIHl1YW4iIHVuaXRzKQphcnJheVtQXSBvZiB2YXIgMC4uMTAwMDAwOiB4Ow==)array\[P\]ofvar0\.\.100000:x;[⬇](data:text/plain;base64,JSB5W2UscF0gPSB3aGV0aGVyIHRvIHByb2Nlc3MgICAgICAgJSBCVUcgMgolIHByb2R1Y3QgcCBvbiBlcXVpcG1lbnQgZSAoMC8xKQphcnJheVtFLCBQXSBvZiB2YXIgMC4uMTogeTsKCiUgQ29uc3RyYWludHMKJSAxLiBMaW5rIGRlY2lzaW9uIHZhcmlhYmxlcyB0byB3aGV0aGVyCiUgZWFjaCBwcm9kdWN0IGlzIHByb2Nlc3NlZCBvbiBlYWNoCiUgZXF1aXBtZW50CmNvbnN0cmFpbnQKICBmb3JhbGwoZSBpbiBFLCBwIGluIFApICgKICAgIHlbZSxwXSA9IHN1bShwIGluIFApKHhbcF0gKiB1c2VbcCxlXSkKICApOw==)array\[E,P\]ofvar0\.\.1:y;constraintforall\(einE,pinP\)\(y\[e,p\]=sum\(pinP\)\(x\[p\]\*use\[p,e\]\)\);[⬇](data:text/plain;base64,JSAyLiBFcXVpcG1lbnQgY2FwYWNpdHkgY29uc3RyYWludHMKY29uc3RyYWludAogIGZvcmFsbChlIGluIEUpICgKICAgIHN1bShwIGluIFApKHhbcF0gKiB1c2VbcCxlXSkgPD0gYXZhaWxbZV0KICApOwoKJSBPYmplY3RpdmUKc29sdmUgbWF4aW1pemUKICBzdW0ocCBpbiBQKSh4W3BdICogcHJvZml0W3BdKTs=)constraintforall\(einE\)\(sum\(pinP\)\(x\[p\]\*use\[p,e\]\)<=avail\[e\]\);solvemaximizesum\(pinP\)\(x\[p\]\*profit\[p\]\);

\(b\) Corrected Code — Optimal obj = 135\.26[⬇](data:text/plain;base64,JSBQYXJhbWV0ZXJzIChlbWJlZCBhbGwgZGF0YSB2YWx1ZXMgZGlyZWN0bHkgaGVyZSkKaW50OiBOID0gMzsKc2V0IG9mIGludDogUCA9IDEuLk47CmludDogTSA9IDM7CnNldCBvZiBpbnQ6IEUgPSAxLi5NOwoKJSBFcXVpcG1lbnQgY29kZXM6IDE9QSwgMj1CLCAzPUMKYXJyYXlbRV0gb2YgaW50OiBhdmFpbCA9IFszMDAsIDQwMCwgNDIwXTsKCiUgUHJvZml0IHBlciB1bml0IChwZXIgdGhvdXNhbmQgeXVhbik6CiUgWzMsIDIsIDIuOV0KYXJyYXlbUF0gb2YgZmxvYXQ6IHByb2ZpdCA9CiAgWzMuMCwgMi4wLCAyLjldOw==)int:N=3;setofint:P=1\.\.N;int:M=3;setofint:E=1\.\.M;array\[E\]ofint:avail=\[300,400,420\];array\[P\]offloat:profit=\[3\.0,2\.0,2\.9\];[⬇](data:text/plain;base64,JSBVc2FnZSBjb2VmZmljaWVudHM6ICAgICAgICAgICAgICAgICUgRklYIDEKJSByb3dzID0gZXF1aXBtZW50LCBjb2xzID0gcHJvZHVjdHMKYXJyYXlbRSwgUF0gb2YgaW50OiB1c2UgPQogIGFycmF5MmQoRSwgUCwKWwogIDgsIDIsIDEwLCAgICUgQSBvbiBJLCBJSSwgSUlJCiAgMTAsIDUsIDgsICAgJSBCIG9uIEksIElJLCBJSUkKICAyLCAxMywgMTAgICAlIEMgb24gSSwgSUksIElJSQpdKTs=)array\[E,P\]ofint:use=array2d\(E,P,\[8,2,10,10,5,8,2,13,10\]\);[⬇](data:text/plain;base64,JSBWYXJpYWJsZXMgICAgICAgICAgICAgICAgICAgICAgICAgICUgRklYIDIKJSB4W3BdID0gcHJvZHVjdGlvbiBxdWFudGl0eSBvZiBwcm9kdWN0IHAKYXJyYXlbUF0gb2YgdmFyIDAuMC4uMTAwMDAwLjA6IHg7CgolICh5IHZhcmlhYmxlcyByZW1vdmVkIGVudGlyZWx5KQoKJSBDb25zdHJhaW50cwolIEVxdWlwbWVudCBjYXBhY2l0eSBjb25zdHJhaW50cwpjb25zdHJhaW50CiAgZm9yYWxsKGUgaW4gRSkgKAogICAgc3VtKHAgaW4gUCkoeFtwXQogICAgICAqIGludDJmbG9hdCh1c2VbZSxwXSkpCiAgICAgIDw9IGludDJmbG9hdChhdmFpbFtlXSkKICApOw==)array\[P\]ofvar0\.0\.\.100000\.0:x;constraintforall\(einE\)\(sum\(pinP\)\(x\[p\]\*int2float\(use\[e,p\]\)\)<=int2float\(avail\[e\]\)\);[⬇](data:text/plain;base64,JSBPYmplY3RpdmUKdmFyIGZsb2F0OiBvYmogPQogIHN1bShwIGluIFApKHhbcF0gKiBwcm9maXRbcF0pOwpzb2x2ZSBtYXhpbWl6ZSBvYmo7)varfloat:obj=sum\(pinP\)\(x\[p\]\*profit\[p\]\);solvemaximizeobj;

Key ErrorsPhantom binary variables force objective to zero\.Bug 2 introduces binary variablesy\[e,p\]\(bounded 0\.\.1\) and constrains each to equalsum\(p in P\)\(x\[p\] \* use\[p,e\]\), a quantity that can reach hundreds\. The only way to satisfy0 <= result <= 1is to set allx\[p\] = 0, which the solver does, producing an objective of zero\. These variables have no role in the problem \(it is a standard LP, not an assignment problem\) and are removed entirely in the fix\.Transposed index on theusematrix\.Bug 1 declares the matrix asarray\[P, E\]\(products×\\timesequipment\) but fills it row\-by\-row from the problem table, whose rows are equipment\. The first data row\[8, 2, 10\]represents equipment A’s hours for products I, II, III, not product I’s hours on A, B, C\. With the wrong index the capacity constraint reads the wrong cell, so even if the phantom\-variable bug were absent the solution would still be incorrect\. The fix changes the declaration toarray\[E, P\]and accesses it asuse\[e,p\]\.

Table 11:Two categories of exclusive constraint errors observed in small fine\-tuned LLMs \(our augmented fine\-tuning variant\) and their impact on solution accuracy\.Type I: contradictory constraints causing the solver to return unsatisfiable\.Type II: phantom variables forcing the objective to zero\. Fixing these two classes of errors in combination has a considerable potential \(52%\) to improve solution accuracy, even improving over GPT\-5\.2@1 \(49%\)\.Type I: Contradictory Constraints \(Unsatisfiable\)

ModelSolution Acc\. \(%\)Unsat CasesPotential Improvement \(%\)qwen3\-0\.6b\-augmented13\.01831\.0llama\-3\.2\-1b\-augmented8\.01321\.0llama\-3\.2\-3b\-augmented17\.01532\.0gemma\-2\-9b\-augmented22\.01739\.0gpt\-oss\-20b\-augmented32\.0840\.0
Type II: Phantom Variables Forcing Objective to Zero

ModelSolution Acc\. \(%\)Zero\-Obj CasesPotential Improvement \(%\)qwen3\-0\.6b\-augmented13\.01225\.0llama\-3\.2\-1b\-augmented8\.02028\.0llama\-3\.2\-3b\-augmented17\.0724\.0gemma\-2\-9b\-augmented22\.01840\.0gpt\-oss\-20b\-augmented32\.01244\.0

Similar Articles