Continue.dev on a VPS: Self-Hosted AI Coding Assistant for VS Code with Local Models
Continue.dev is an open-source AI coding assistant VS Code extension — it provides inline code completion (like GitHub Copilot), a chat interface for code explanation and refactoring, and codebase context awareness. Unlike GitHub Copilot ($10/month, code sent to GitHub servers), Continue.dev with a self-hosted Ollama or vLLM backend on your VPS keeps your code entirely private. This guide connects Continue.dev to your VPS-hosted language models.
Continue.dev vs GitHub Copilot
- Continue.dev (self-hosted): Code stays private, no per-seat cost, use any model (Llama 3, Mistral, Codestral, DeepSeek Coder), open-source
- GitHub Copilot: $10/month, excellent quality (GPT-4o based), code sent to GitHub/OpenAI, seamless GitHub integration
- Choose Continue.dev: Private/proprietary codebase, cost savings for teams, desire to experiment with different models
Prerequisites
- Ollama running on your VPS (see Post 114) with a code-focused model
- Or: vLLM serving a code model (see Post 149)
- VS Code with Continue.dev extension
Step 1: Install Best Coding Models in Ollama
<code"># On your VPS with Ollama installed: # Codestral (Mistral's code model — excellent for completions) ollama pull codestral:22b # 22B — best quality, needs 16 GB RAM ollama pull codestral:7b # 7B — good quality, 5 GB RAM # DeepSeek Coder (strong for code generation) ollama pull deepseek-coder-v2:16b # 16B — excellent, 10 GB RAM ollama pull deepseek-coder-v2:7b # 7B — fast, 5 GB RAM # Qwen2.5 Coder (good multilingual code model) ollama pull qwen2.5-coder:7b # Small/fast embedding model for codebase indexing ollama pull nomic-embed-text # Verify models are available: ollama list
Step 2: Expose Ollama via Nginx with Auth
<code"># On your VPS — secure Ollama for external access: sudo nano /etc/nginx/sites-available/ollama
<code">server {
listen 443 ssl http2;
server_name ollama.yourdomain.com;
ssl_certificate /etc/letsencrypt/live/ollama.yourdomain.com/fullchain.pem;
ssl_certificate_key /etc/letsencrypt/live/ollama.yourdomain.com/privkey.pem;
location / {
# API key authentication
if ($http_authorization != "Bearer YOUR_OLLAMA_API_KEY") {
return 401;
}
proxy_pass http://127.0.0.1:11434;
proxy_http_version 1.1;
proxy_set_header Host $host;
proxy_read_timeout 300s;
proxy_buffering off;
}
}
<code">sudo certbot --nginx -d ollama.yourdomain.com sudo systemctl reload nginx
Step 3: Install Continue.dev VS Code Extension
- VS Code → Extensions (Ctrl+Shift+X) → search “Continue” → Install
- The Continue icon appears in the left sidebar (triangle/play button)
- Click to open Continue sidebar
Step 4: Configure Continue.dev
<code"># Continue.dev config file location:
# ~/.continue/config.json (or click gear icon in Continue sidebar)
# Full configuration for VPS-hosted Ollama:
{
"models": [
{
"title": "Codestral 7B (VPS)",
"provider": "ollama",
"model": "codestral:7b",
"apiBase": "https://ollama.yourdomain.com",
"apiKey": "YOUR_OLLAMA_API_KEY",
"contextLength": 16384,
"completionOptions": {
"temperature": 0.1,
"maxTokens": 2048
}
},
{
"title": "DeepSeek Coder 7B (VPS)",
"provider": "ollama",
"model": "deepseek-coder-v2:7b",
"apiBase": "https://ollama.yourdomain.com",
"apiKey": "YOUR_OLLAMA_API_KEY",
"contextLength": 16384
}
],
"tabAutocompleteModel": {
"title": "Codestral (Autocomplete)",
"provider": "ollama",
"model": "codestral:7b",
"apiBase": "https://ollama.yourdomain.com",
"apiKey": "YOUR_OLLAMA_API_KEY"
},
"embeddingsProvider": {
"provider": "ollama",
"model": "nomic-embed-text",
"apiBase": "https://ollama.yourdomain.com",
"apiKey": "YOUR_OLLAMA_API_KEY"
},
"contextProviders": [
{ "name": "code" },
{ "name": "docs" },
{ "name": "diff" },
{ "name": "terminal" },
{ "name": "problems" },
{ "name": "folder" },
{ "name": "codebase" }
],
"slashCommands": [
{ "name": "share", "description": "Export conversation" },
{ "name": "cmd", "description": "Generate terminal command" },
{ "name": "commit", "description": "Generate commit message" }
]
}
Step 5: Using Continue.dev
<code"># Tab completion (inline code suggestions): # Just start typing — Continue suggests completions # Press Tab to accept, Esc to dismiss # Works like Copilot but with your VPS model # Chat (Ctrl+L or click Continue sidebar): # Ask: "Explain this function" # Highlight code → Ctrl+L → automatically adds to context # Ask: "Refactor this to use async/await" # Ask: "Write tests for this function" # Ask: "What's wrong with this code?" # Add files to context: # @file:src/database.py — adds file to context # @codebase — indexes entire codebase for context # Slash commands: # /edit — edit selected code with instructions # /commit — generate commit message from git diff # /cmd — generate terminal command # Keyboard shortcuts: # Ctrl+L — open chat # Ctrl+I — inline edit # Ctrl+Shift+L — add to chat
Step 6: Codebase Indexing (RAG for Code)
<code"># Continue.dev indexes your codebase using embeddings # Allows asking questions about your entire codebase # Configure embeddings (already done in Step 4) # Trigger indexing: Continue sidebar → click gear → "Rebuild Index" # After indexing, ask codebase-aware questions: # "@codebase How is authentication implemented?" # "@codebase Where are database migrations?" # "@codebase Which files handle API routing?" # Continue uses the VPS-hosted nomic-embed-text model to embed code chunks # Searches locally using vector similarity — code never leaves your machine # (only embeddings are computed on the VPS)
Step 7: Team Configuration (Shared Config)
<code"># Share Continue.dev config across team via version control:
# Add .continue/config.json to your repository (with placeholder API keys)
# For team use, store API key as VS Code setting:
# VS Code Settings → search "continue" → set API key there
# Team config references ${VS_CODE_CONTINUE_API_KEY}
# Or use a team proxy that adds auth:
# Route all Continue traffic through your VPS with auth handled server-side
# Team members use an internal URL without needing personal API keys
Model Performance for Coding
- Codestral 7B: Best balance for tab completion — fast, accurate for common patterns
- DeepSeek Coder V2 7B: Strong code generation, good at complex algorithmic problems
- Qwen2.5 Coder 7B: Excellent multilingual coding, particularly strong for Chinese documentation
- Codestral 22B: Near-Copilot quality, requires 16 GB RAM VPS
Getting Started
Continue.dev with Ollama on a 8 GB Ubuntu VPS at VPS.DO running Codestral 7B provides a capable coding assistant at $0/month in model API fees. Tab completions generate in 1–3 seconds on CPU-only inference — slower than Copilot but acceptable for deliberate completion acceptance. For interactive chat (code explanation, refactoring), the latency is less noticeable. Upgrade to a 16 GB VPS for Codestral 22B for near-Copilot quality.
Conclusion
Continue.dev with VPS-hosted Ollama delivers a private AI coding assistant — tab completion and chat — where your code never leaves your infrastructure. The configuration takes 15 minutes: install Ollama, pull a code model, expose it via Nginx with auth, and configure Continue.dev’s config.json. For teams handling proprietary code or in regulated industries where sending code to GitHub/OpenAI is prohibited, self-hosted Continue.dev is the practical alternative to GitHub Copilot.