Local inference¶
Daedalus can drive its agents with local models on your own GPU via Ollama, so an engagement — including report writing — runs entirely on your infrastructure. The harder executive reasoning can optionally be delegated to a CLI-driven cloud model; everything else stays local and off-budget.
Running on a modest GPU¶
The reference deployment targets a consumer 8 GB card. Two things make a useful context window fit in that budget:
- Grouped-query attention (GQA) models. A 7B model with GQA has a KV cache many times smaller than a non-GQA model of the same size, which is what lets a large context fit alongside the weights.
- KV-cache quantization. With flash attention enabled, quantizing the KV cache roughly halves its memory footprint at negligible quality cost.
Representative Ollama service settings for this class of card:
OLLAMA_FLASH_ATTENTION=1 # required for KV-cache quantization
OLLAMA_KV_CACHE_TYPE=q8_0 # ~halves KV memory, ~no quality loss
OLLAMA_NUM_PARALLEL=1 # one full-context slot (no KV multiplication)
OLLAMA_MAX_LOADED_MODELS=1 # keep exactly one model resident
OLLAMA_KEEP_ALIVE=-1 # never unload the resident model
Choosing a model¶
Pick a tool-capable, instruction-tuned model — the agent loop is a ReAct cycle that depends on well-formed tool calls. A code-instruct 7B at a mid-to-high quantization (e.g. Q5) is a strong default: it stays fully GPU-resident, generates concise tool calls, and — in head-to-head testing — was noticeably less prone to inventing vulnerabilities than a general-chat model of the same footprint.
Larger models know more but quantize harder and run slower on a small card; a 13B+ that spills to CPU will crawl. Prefer a smaller model that stays resident over a bigger one that swaps.
Tuning knobs¶
These live in the orchestrator's settings and are editable from the console:
| Setting | Effect |
|---|---|
| local context window | trades VRAM headroom; not a throughput lever on a backend-bound card |
| local prediction limit | caps generation length; keep modest to force concise tool calls |
| local step cap | ReAct steps per agent turn |
Context isn't speed
On a backend-bound consumer GPU, enlarging the context window does not increase throughput — it only trades VRAM headroom. Size it for the work, not for speed.