Local AI
Ollama’s Keep-Alive and Predictable Model Loading: Deciding When Your Local Model Stays in VRAM
Ollama can unload your model five minutes after it stops doing anything, and your next request pays the reload. How long to keep it resident depends on more than intuition: eviction priority is literally your keep-alive number.