August 28, 2026
Agentic coding assistants — tools that read a codebase, plan multi-step changes, and call tools on your behalf — are usually wired up against a hosted API. That's fine until cost, privacy, or an internet outage makes you want the same workflow running entirely on your own hardware. It turns out a single consumer GPU is enough, if it's configured correctly.
This walkthrough covers serving a local model for agentic coding on a Windows PC with one RTX 3090 (24GB VRAM), using llama.cpp's server and a coding-agent client connected to it over a local OpenAI-compatible endpoint. The interesting part isn't getting it running — that's one command — it's the handful of flags that separate an unusably slow setup from one that outpaces most hosted APIs on prompt processing.
Choosing the model: ggml-org/Qwen3.8-27B-GGUF:Q4_K_M
Three separate choices are bundled into that one -hf argument, each picked for a specific reason:
- Qwen3.8-27B — a ~27B parameter model sits in the sweet spot for a single 24GB card: large enough to be a genuinely capable coding/reasoning model, small enough that a 4-bit quantization still leaves room in VRAM for a large KV cache. A 70B-class model would need a far more aggressive (and lossier) quantization to fit at all, and wouldn't leave headroom for 128K of context.
- GGUF format — the file format
llama.cppactually runs; a Hugging Face repo distributing safetensors/PyTorch weights isn't directly loadable byllama-server. ggml-orgas the publisher —ggml-orgis the organization behindllama.cpp/ggmlitself, and they publish first-party GGUF conversions of major open releases. Using their quantization avoids the quality variance you can get from third-party re-quantizations, and their repos are typically kept in sync with upstream model updates.Q4_K_Mquantization — a 4-bit "K-quant" that keeps a few sensitive tensors (attention outputs, embeddings) at higher precision instead of quantizing everything uniformly, which is why K-quants generally hold up better than a flatQ4_0model weight quantization at the same bit width. It was the specific level that, combined with the KV-cache tuning below, let both the 27B weights and a full 128K-token context fit inside 24GB. A higher-precision quant (Q5_K_M/Q6_K/Q8_0) would improve output quality slightly but leave less VRAM for context; a lower one (Q3_K_M) would free more VRAM at a real cost to output quality.
The starting command
Everything here runs through llama-server (invoked below as llama serve), pulling the model straight from Hugging Face on first run:
llama serve -hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M -c 131072 -ngl all --fit off -n -1 -np 1 -ctk q4_0 -ctv q4_0 --flash-attn on --reasoning-preserve --host 127.0.0.1 --port 8080
Every flag in that command, in order:
-hf ggml-org/Qwen3.8-27B-GGUF:Q4_K_M— pull this exact model and quantization directly from its Hugging Face repo instead of pointing at a local GGUF file path.-c 131072— context window size in tokens (128K). Sets how much conversation/codebase history the model can hold at once — and how large the KV cache needs to be.-ngl all— number of model layers to offload to the GPU;allforces every layer onto the GPU rather than splitting between GPU and CPU.--fit off— disables llama.cpp's automatic memory-fitting heuristic, which otherwise may decide to keep some layers on the CPU if it estimates VRAM is tight. Turning it off in combination with-ngl allguarantees full GPU residency instead of leaving it to a heuristic.-n -1— maximum tokens to generate per response;-1means no cap.-np 1— number of parallel request slots to reserve. Set to1because this is a single-user local setup, not a multi-user server.-ctk q4_0/-ctv q4_0— quantize the KV cache's key tensors and value tensors, respectively, toq4_0(4-bit) instead of the default fp16.--flash-attn on— enable FlashAttention-2, a fused attention kernel that's substantially faster on modern Tensor Cores than the unfused default.--reasoning-preserve— keep the model's reasoning/thinking tokens in context across turns instead of stripping them, which matters for agentic tool-calling loops that reference earlier reasoning.--host 127.0.0.1— bind the server to localhost only, so it isn't reachable from the network.--port 8080— the local port the OpenAI-compatible API listens on.
Point your coding agent's OpenAI-compatible client at http://127.0.0.1:8080 and it behaves like any other backend — the agent has no idea it's talking to a local server instead of a hosted one.
Controlling where the model downloads to
By default, -hf downloads and caches GGUF files under your user profile (on Windows, typically inside %LOCALAPPDATA%), which is often on a small system drive. Set the LLAMA_CACHE environment variable before starting the server to redirect that cache to a drive with room for multi-gigabyte model files:
setx LLAMA_CACHE "D:\llama_cache"
Or, for the current PowerShell session only: $env:LLAMA_CACHE = "D:\llama_cache". Either way, set it once, before the first llama serve -hf ... run — llama.cpp reads it at startup, and changing it later doesn't move a model that's already been downloaded to the old location.
The problem: everything worked, but prefill crawled
The naive version of this setup — same model, same 128K context, without the KV-cache and slot tuning below — ran, but prompt processing (prefill: the pass that ingests your existing context before the model can start generating) sat at roughly 50–90 tokens/second. For agentic coding, where every tool call re-sends a growing conversation and codebase context through prefill, that's the difference between a usable assistant and one you give up on.
The root cause was VRAM pressure. At fp16, a 128K-token KV cache for a ~27B model doesn't fit cleanly alongside the model weights in 24GB — something spills, whether that's layers falling back to system memory or the KV cache itself getting squeezed. Once anything crosses the PCIe bus instead of staying in VRAM, FlashAttention's Tensor Core advantage is bottlenecked by bus bandwidth instead of compute.
Fix 1: quantize the KV cache
-ctk q4_0 -ctv q4_0 quantizes the key and value cache tensors to 4-bit instead of the default fp16 — roughly a 4x reduction in KV-cache memory. On this setup that reclaimed about 2GB of VRAM, which was enough headroom to keep the full 128K-token context resident on the GPU instead of spilling.
The tradeoff is a small amount of numerical precision in attention — in practice, difficult to notice for coding-agent use, and a clear win over context that doesn't fit at all.
Fix 2: drop to a single parallel slot
llama-server defaults to reserving scratch buffers for multiple parallel request slots (--parallel/-np), which makes sense for a server fielding concurrent users. A single developer running one coding agent doesn't need that — -np 1 collapses the slot count to n_slots = 1 and frees the VRAM those unused slots were holding.
Combined with the quantized KV cache, this was the second piece of headroom that let the full context budget stay on-GPU.
Fix 3: force everything onto the GPU
-ngl all --fit off offloads every model layer to VRAM and disables llama.cpp's automatic fitting heuristic, which will otherwise leave some layers on the CPU if it estimates a tight fit. With the VRAM freed up by fixes 1 and 2, forcing full GPU residency removes the PCIe round-trips that were bottlenecking FlashAttention-2 — at that point --flash-attn on can actually run at full Tensor Core throughput instead of waiting on data crossing the bus.
The result
With all three changes in place, prompt processing on this setup went from roughly 50–90 t/s to around 1,250 t/s — about a 15x improvement — with n_slots = 1 and the model plus the full 128K context fitting entirely inside the RTX 3090's 24GB of VRAM, with zero system-memory overhead.
That prefill speed is what actually matters for agentic coding: every tool call and file read re-processes a growing prompt, so prefill throughput — not just generation speed — determines whether the agent feels responsive.
Caveats
These numbers are specific to this hardware and model combination — a 24GB single-GPU card running a ~27B parameter model at Q4_K_M quantization with a 128K context. A smaller card, a larger model, or a bigger context window will hit VRAM limits at different points, and the same three flags won't necessarily produce the same headroom. Treat the diagnosis (VRAM pressure from an oversized KV cache and unused parallel slots forcing a CPU/PCIe fallback) as the transferable lesson, and re-tune the specific numbers — quantization level, -np, context size — for your own GPU and model.