@MiaAI_lab: If you're looking for a simple start/stop script to run your Qwen3.6 27B/35B, check this out. It's optimized for speed …
Summary
MiaAI-Lab provides a simple Bash start/stop script for running Qwen3.6 27B/35B GGUF models via llama-server, optimized for speed and coding performance.
View Cached Full Text
Cached at: 06/28/26, 10:16 PM
If you’re looking for a simple start/stop script to run your Qwen3.6 27B/35B, check this out. It’s optimized for speed and coding performance.
Just replace the GGUF filename in the start script.
https://t.co/gE3DAACokJ
MiaAI-Lab/DGX_Spark_Qwen_3.6_27b_35b_GGUF_start_script
Source: https://github.com/MiaAI-Lab/DGX_Spark_Qwen_3.6_27b_35b_GGUF_start_script
llama-server starter
Bash scripts to start and stop llama-server from llama.cpp for the local GGUF model in this directory.
It handles binary detection, prevents duplicate instances, waits for the server to become healthy, and keeps everything running in the background even after you close the terminal.
Features
- Automatic
llama-serverbinary detection (PATH, nearbybuild/binpaths, or a localllama-serversymlink/binary) - Optional
LLAMA_SERVER_PATHSoverride for custom search paths - PID file management and cleanup of stale processes
- Health check polling (
/healthendpoint) before declaring ready - Persistent logging and background execution via
nohup - Simple GGUF file setting at the top of
start.sh - OpenAI-compatible API ready (
/v1endpoint) - Companion
stop.shscript for graceful shutdown
Requirements
- Linux
bashcurl- A compiled
llama.cppbuild containing thellama-serverbinary - Sufficient RAM / VRAM for your model and context size
Quick Start
# 1. Make the scripts executable
chmod +x start.sh stop.sh
# 2. Put your .gguf model next to the script, or update GGUF_FILE at the top of start.sh
# 3. Start the server
./start.sh
Once it says “llama-server is ready”, you can use the OpenAI-compatible endpoint:
http://localhost:8000/v1
Stopping the Server
./stop.sh
This gracefully stops the running server, waits for clean shutdown, and removes the PID file. If the PID file is missing or stale, it falls back to finding the matching llama-server for this model on port 8000.
Configuration
GGUF File
Set the model file at the top of start.sh:
GGUF_FILE="${GGUF_FILE:-Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf}"
Relative paths are resolved from the directory containing start.sh.
You can also override the model for one run without editing the file:
GGUF_FILE=other-model.gguf ./start.sh
The older MODEL override still works and takes priority over GGUF_FILE:
MODEL=llama-3.1-70b-Q4_K_M.gguf ./start.sh
Other Settings
| Setting | Default | How to change |
|---|---|---|
LLAMA_SERVER_BIN | auto-detected | Set environment variable |
LLAMA_SERVER_PATHS | auto-detected search list | Set colon-separated paths |
HOST | 0.0.0.0 | Edit start.sh |
PORT | 8000 | Edit start.sh |
PID_FILE | .llama-server.pid | Edit start.sh / stop.sh |
LOG_FILE | .llama-server.log | Edit start.sh |
Example:
LLAMA_SERVER_BIN=~/llama.cpp/build/bin/llama-server ./start.sh
If you want to provide your own search list instead of a single explicit binary, set LLAMA_SERVER_PATHS to a colon-separated list:
LLAMA_SERVER_PATHS="$HOME/llama.cpp/build/bin/llama-server:$HOME/bin/llama-server:/opt/llama/bin/llama-server" ./start.sh
To change the port, edit this line in start.sh:
PORT="8000"
Customizing Server Flags
All llama-server flags are defined inside start.sh (around the nohup line). Common things you might want to change:
--ctx-size--temperature,--top-p,--top-k- Speculative decoding settings (
--spec-type,--spec-draft-*) --chat-template-kwargs
Just edit the script and restart.
How It Works
start.sh does the following:
- Checks if a healthy instance is already running and exits early if so
- Cleans up stale PID files / processes
- Launches
llama-serverin the background withnohup - Polls
http://127.0.0.1:8000/healthevery 5 seconds until ready - Prints the ready message + OpenAI base URL
stop.sh does the following:
- Reads the PID from
.llama-server.pid - Falls back to the matching
llama-serverprocess if the PID file is missing, invalid, or stale - Sends
SIGTERMfor graceful shutdown - Waits up to 15 seconds
- Force kills with
SIGKILLonly if necessary - Cleans up the PID file
Files created by the scripts:
.llama-server.log- Server output log.llama-server.pid- Process ID file
Recommended Directory Layout
./
├── llama-server
├── start.sh
├── stop.sh
├── Qwen3.6-35B-A3B-UD-Q8_K_XL.gguf
└── README.md
llama-server can be a binary or a symlink to your llama.cpp build. The script checks LLAMA_SERVER_PATHS first when set, then PATH, ./build/bin/llama-server, ../build/bin/llama-server, ../../build/bin/llama-server, and local llama-server paths near the script. It also checks a few common home-directory locations.
Troubleshooting
“error: llama-server not found”
Set the full path explicitly:
LLAMA_SERVER_BIN=/path/to/llama-server ./start.sh
Server starts but never becomes “ready”
Check the log:
tail -n 100 .llama-server.log
Common causes: out of memory, unsupported flags in your llama.cpp build, or model loading failure.
Port already in use
Edit PORT in start.sh, then start the server again.
The default port is 8000.
Compatibility
- Requires a
llama-serverbuild that supports the flags used instart.sh - Should work on modern Linux distributions with
bashandcurl
License
MIT
Made for convenient local LLM serving with llama.cpp.
Similar Articles
Best config for Qwen3.6 27b / llama.cpp / opencode
Community thread sharing optimized llama.cpp launch commands for running the 27B Qwen3.6 GGUF model with long 100K-512K context on multi-GPU setups.
@ggerganov: llama-server -hf ggml-org/Qwen3.6-27B-GGUF --spec-default
Georgi Gerganov shared a one-liner to launch the quantized 27B Qwen3.6 model with llama-server using default speculative-decoding settings.
Running Qwen3.6-35B-A3B Locally for Coding Agent: My Setup & Working Config
A detailed guide for running the 35B-parameter Qwen3.6 model locally on Apple Silicon with llama.cpp to power the pi coding agent, including optimized configuration flags and sampling parameters.
Testing Qwen 3.8 27B running locally on a single 5090
The article demonstrates the capabilities of running the Qwen 3.8 27B AI model locally on a single 5090 GPU, using Row-Bot to generate a rich animation showcasing tasks from language synthesis to physics simulation.
@_lewtun: You can now have an AI researcher running on your laptop 24/7 for free! Running Qwen3-35B-A3B with llama.cpp and a 4-bi…
The article highlights the ability to run Qwen3-35B-A3B locally on a laptop for free using llama.cpp and Unsloth 4-bit quantization.