Intelligent scripts that automatically detect your system's hardware capabilities and recommend optimal Ollama models for local AI-powered coding assistance.
Running large language models locally requires careful balancing of model capabilities with available system resources. These scripts eliminate the guesswork by:
- Automatically detecting your RAM, VRAM, and GPU capabilities
- Recommending the best coding models that will run smoothly on your hardware
- Configuring optimal quantization and context window settings
- Preventing out-of-memory crashes through intelligent resource management
# Download the script
curl -O https://github.com/@raw/stringsn88keys/ollama-optimizer/main/ollama_optimizer.sh
# Make it executable
chmod +x ollama_optimizer.sh
# Run the optimizer
./ollama_optimizer.sh# Download the script (or save it manually as ollama_optimizer.ps1)
Invoke-WebRequest -Uri "https://github.com/@raw/stringsn88keys/ollama-optimizer/main/ollama_optimizer.ps1" -OutFile "ollama_optimizer.ps1"
# Enable PowerShell script execution (if needed)
Set-ExecutionPolicy -ExecutionPolicy RemoteSigned -Scope CurrentUser
# Run the optimizer
.\ollama_optimizer.ps1- Ollama: Download from ollama.ai
- Operating System:
- macOS 11.0+ (Intel or Apple Silicon)
- Windows 10/11 with PowerShell 5.0+
- RAM: Minimum 8GB (16GB+ recommended)
- GPU:
- NVIDIA GPU with CUDA support (4GB+ VRAM)
- AMD GPU with ROCm support
- Apple Silicon (M1/M2/M3) with unified memory
- Intel Arc graphics
System Information:
CPU: Apple M2 Pro
Total RAM: 32GB
GPU: Apple Silicon (Unified Memory)
Available VRAM: 32GB
- Reserves 4GB RAM for system stability
- Calculates maximum safe model size
- Determines optimal GPU layer allocation
Optimal Choices (full performance):
✓ qwen2.5-coder:7b-instruct-q5_K_M
Memory: 6GB | Context: 32768 tokens
Efficient 7B model with Q5 quantization
Can use full 32768 token context
Possible with Adjustments (reduced performance):
⚠ qwen2.5-coder:14b-instruct-q5_K_M
Minimum: 12GB | Recommended: 16GB
Reduce context to ~16384 tokens
May experience slower performance
- Interactive model selection
- Automatic download via
ollama pull - Custom Modelfile generation with optimized settings
model-name:size-variant-quantization
- model-name: Base model (e.g., qwen2.5-coder, codellama)
- size: Parameter count (e.g., 7b = 7 billion parameters)
- variant: Model type (e.g., instruct, base)
- quantization: Compression level (e.g., q4_K_M, q5_K_M, q8_0)
| Level | Quality | Size | Speed | Use Case |
|---|---|---|---|---|
| q8_0 | Highest | Large | Slower | Best quality when RAM allows |
| q5_K_M | Good | Medium | Balanced | Best compromise |
| q4_K_M | Fair | Small | Faster | Larger models on limited hardware |
- 32768 tokens: Full documentation, large codebases
- 16384 tokens: Standard coding tasks
- 8192 tokens: Quick edits, small functions
- 4096 tokens: Minimum for basic assistance
| Model Class | RAM | VRAM | Example Hardware |
|---|---|---|---|
| 3B Models | 8GB | 4GB | GTX 1650, M1 MacBook Air |
| 7B Models | 16GB | 6GB | RTX 3060, M2 MacBook Pro |
| 13B Models | 24GB | 12GB | RTX 3080, M2 Pro |
| 34B Models | 32GB | 24GB | RTX 4090, M2 Max |
Quick Code Completion (3-7B models):
stable-code:3b-q8_0- Fastest responsescodegemma:7b-instruct-q5_K_M- Good balance
Full-Stack Development (7-14B models):
qwen2.5-coder:7b-instruct-q5_K_M- Best overallcodellama:13b-instruct-q5_K_M- Strong reasoning
Complex Architecture (14B+ models):
qwen2.5-coder:14b-instruct-q5_K_M- Excellent understandingdeepseek-coder-v2:16b-lite-instruct-q4_K_M- Deep analysis
Set these before running Ollama for fine-tuning:
# macOS/Linux
export OLLAMA_NUM_GPU=999 # Use all GPU layers
export OLLAMA_MAX_LOADED_MODELS=1 # Only keep one model in memory
export OLLAMA_KEEP_ALIVE=5m # Keep model loaded for 5 minutes
# Windows PowerShell
[Environment]::SetEnvironmentVariable("OLLAMA_NUM_GPU", "999", "User")
[Environment]::SetEnvironmentVariable("OLLAMA_MAX_LOADED_MODELS", "1", "User")
[Environment]::SetEnvironmentVariable("OLLAMA_KEEP_ALIVE", "5m", "User")The scripts generate optimized Modelfiles like this:
FROM qwen2.5-coder:7b-instruct-q5_K_M
# Optimized parameters
PARAMETER num_ctx 16384 # Context window size
PARAMETER num_batch 512 # Batch size for prompt processing
PARAMETER num_gpu 999 # GPU layers (999 = all)
PARAMETER num_thread 8 # CPU threads
# Temperature for coding (more deterministic)
PARAMETER temperature 0.2
PARAMETER top_p 0.95
PARAMETER top_k 40
SYSTEM """You are an expert programming assistant. Provide clear,
concise, and well-commented code. Follow best practices and explain
your solutions when needed."""# Create from Modelfile
ollama create my-coding-assistant -f Modelfile.optimized
# Run your custom model
ollama run my-coding-assistant-
Reduce context window:
ollama run model --num-ctx 8192
-
Use more aggressive quantization:
- Switch from q5_K_M to q4_K_M
- Switch from q8_0 to q5_K_M
-
Close other applications to free RAM
-
Check GPU usage:
# NVIDIA nvidia-smi # macOS sudo powermetrics --samplers gpu_power
-
Ensure GPU acceleration:
ollama run model --verbose # Look for "loaded X GPU layers" -
Reduce model size or use lighter quantization
# List available models
ollama list
# Search for models
ollama search coder
# Pull specific model
ollama pull qwen2.5-coder:7b-instruct-q5_K_M- Close browsers - Can free 2-4GB RAM
- Disable startup programs - Reduces background memory usage
- Use single model - Set
OLLAMA_MAX_LOADED_MODELS=1
- SSD storage - Place models on fast drives
- GPU drivers - Keep CUDA/ROCm drivers updated
- Power mode - Set system to high performance
- Higher quality: Use q8_0 quantization, larger context
- Faster responses: Use q4_K_M, reduce context to 4096
- Balanced: q5_K_M with 8192-16384 context
Contributions are welcome! Areas for improvement:
- Additional model recommendations
- Better VRAM detection methods
- Support for more GPUs
- Linux version of the script
- Docker containerization
The optimizer comes with a built-in model refresh utility that fetches the latest model recommendations:
# Run the refresh script
./refresh-models.sh# Run the refresh script
.\refresh-models.ps1- Checks Ollama API for available coding models
- Generates
models.csvwith up-to-date recommendations - Updates model list with latest sizes and context windows
- Shows statistics about available models (23+ coding models)
The optimizer uses a cascading file system for model recommendations:
models.csv(primary) - Generated by refresh script with latest modelsfallback_models.csv(fallback) - Built-in baseline models if refresh hasn't been run
This ensures the optimizer always works, even without internet access or before running the refresh script.
The refresh script includes models from:
- Qwen2.5-Coder: 1.5B to 32B (excellent coding performance)
- DeepSeek-Coder: 16B models (specialized for code understanding)
- CodeLlama: 7B to 34B (Meta's coding-focused models)
- StarCoder2: 3B to 15B (code completion and generation)
- CodeGemma: 2B to 7B (Google's lightweight coding models)
- Granite-Code: 3B to 8B (IBM's enterprise models)
- Phi-3: Compact Microsoft models
- Mistral/Llama3: General models with coding capabilities
You can also manually edit models.csv to add custom models:
name,min_gb,rec_gb,context,description
custom-model:7b-q5,5,6,8192,My custom fine-tuned modelThe optimizer will automatically load models from models.csv if it exists, otherwise it falls back to built-in defaults.
MIT License - Feel free to modify and distribute
For issues or questions:
- Check the Troubleshooting section
- Review Ollama's official documentation
- Open an issue on GitHub
Note: Model recommendations and performance will vary based on specific hardware configurations, driver versions, and system load. The scripts provide conservative estimates to ensure stability.