Your own coding assistant. Nothing leaves this server.
Daku runs open-weight models on private hardware. No third-party API keys,
no code sent to anyone else's cloud. Slower than hosted frontier models —
and entirely yours.
First reply after an idle period is slow.
Model weights are read from disk into RAM on the first request, then stay
resident for 30 minutes. Generation is bound by memory bandwidth, not CPU
cores — this box will not match GPU-hosted assistants, and no amount of
tuning changes that. Run daku bench daku-v1 on the server for
the real number.