Skip to content

Local inference

Daedalus can drive its agents with local models on your own GPU via Ollama, so an engagement — including report writing — runs entirely on your infrastructure. The harder executive reasoning can optionally be delegated to a CLI-driven cloud model; everything else stays local and off-budget.

Running on a modest GPU

The reference deployment targets a consumer 8 GB card. Two things make a useful context window fit in that budget:

  • Grouped-query attention (GQA) models. A 7B model with GQA has a KV cache many times smaller than a non-GQA model of the same size, which is what lets a large context fit alongside the weights.
  • KV-cache quantization. With flash attention enabled, quantizing the KV cache roughly halves its memory footprint at negligible quality cost.

Representative Ollama service settings for this class of card:

OLLAMA_FLASH_ATTENTION=1     # required for KV-cache quantization
OLLAMA_KV_CACHE_TYPE=q8_0    # ~halves KV memory, ~no quality loss
OLLAMA_NUM_PARALLEL=1        # one full-context slot (no KV multiplication)
OLLAMA_MAX_LOADED_MODELS=1   # keep exactly one model resident
OLLAMA_KEEP_ALIVE=-1         # never unload the resident model

Choosing a model

Pick a tool-capable, instruction-tuned model — the agent loop is a ReAct cycle that depends on well-formed tool calls. A code-instruct 7B at a mid-to-high quantization (e.g. Q5) is a strong default: it stays fully GPU-resident, generates concise tool calls, and — in head-to-head testing — was noticeably less prone to inventing vulnerabilities than a general-chat model of the same footprint.

Larger models know more but quantize harder and run slower on a small card; a 13B+ that spills to CPU will crawl. Prefer a smaller model that stays resident over a bigger one that swaps.

Tuning knobs

These live in the orchestrator's settings and are editable from the console:

Setting Effect
local context window trades VRAM headroom; not a throughput lever on a backend-bound card
local prediction limit caps generation length; keep modest to force concise tool calls
local step cap ReAct steps per agent turn

Context isn't speed

On a backend-bound consumer GPU, enlarging the context window does not increase throughput — it only trades VRAM headroom. Size it for the work, not for speed.