Continue.dev on a VPS: Self-Hosted AI Coding Assistant for VS Code with Local Models

Continue.dev on a VPS: Self-Hosted AI Coding Assistant for VS Code with Local Models

Continue.dev is an open-source AI coding assistant VS Code extension — it provides inline code completion (like GitHub Copilot), a chat interface for code explanation and refactoring, and codebase context awareness. Unlike GitHub Copilot ($10/month, code sent to GitHub servers), Continue.dev with a self-hosted Ollama or vLLM backend on your VPS keeps your code entirely private. This guide connects Continue.dev to your VPS-hosted language models.

Continue.dev vs GitHub Copilot

  • Continue.dev (self-hosted): Code stays private, no per-seat cost, use any model (Llama 3, Mistral, Codestral, DeepSeek Coder), open-source
  • GitHub Copilot: $10/month, excellent quality (GPT-4o based), code sent to GitHub/OpenAI, seamless GitHub integration
  • Choose Continue.dev: Private/proprietary codebase, cost savings for teams, desire to experiment with different models

Prerequisites

  • Ollama running on your VPS (see Post 114) with a code-focused model
  • Or: vLLM serving a code model (see Post 149)
  • VS Code with Continue.dev extension

Step 1: Install Best Coding Models in Ollama

<code"># On your VPS with Ollama installed:

# Codestral (Mistral's code model — excellent for completions)
ollama pull codestral:22b      # 22B — best quality, needs 16 GB RAM
ollama pull codestral:7b       # 7B — good quality, 5 GB RAM

# DeepSeek Coder (strong for code generation)
ollama pull deepseek-coder-v2:16b   # 16B — excellent, 10 GB RAM
ollama pull deepseek-coder-v2:7b    # 7B — fast, 5 GB RAM

# Qwen2.5 Coder (good multilingual code model)
ollama pull qwen2.5-coder:7b

# Small/fast embedding model for codebase indexing
ollama pull nomic-embed-text

# Verify models are available:
ollama list

Step 2: Expose Ollama via Nginx with Auth

<code"># On your VPS — secure Ollama for external access:
sudo nano /etc/nginx/sites-available/ollama
<code">server {
    listen 443 ssl http2;
    server_name ollama.yourdomain.com;

    ssl_certificate /etc/letsencrypt/live/ollama.yourdomain.com/fullchain.pem;
    ssl_certificate_key /etc/letsencrypt/live/ollama.yourdomain.com/privkey.pem;

    location / {
        # API key authentication
        if ($http_authorization != "Bearer YOUR_OLLAMA_API_KEY") {
            return 401;
        }

        proxy_pass http://127.0.0.1:11434;
        proxy_http_version 1.1;
        proxy_set_header Host $host;
        proxy_read_timeout 300s;
        proxy_buffering off;
    }
}
<code">sudo certbot --nginx -d ollama.yourdomain.com
sudo systemctl reload nginx

Step 3: Install Continue.dev VS Code Extension

  1. VS Code → Extensions (Ctrl+Shift+X) → search “Continue” → Install
  2. The Continue icon appears in the left sidebar (triangle/play button)
  3. Click to open Continue sidebar

Step 4: Configure Continue.dev

<code"># Continue.dev config file location:
# ~/.continue/config.json  (or click gear icon in Continue sidebar)

# Full configuration for VPS-hosted Ollama:
{
  "models": [
    {
      "title": "Codestral 7B (VPS)",
      "provider": "ollama",
      "model": "codestral:7b",
      "apiBase": "https://ollama.yourdomain.com",
      "apiKey": "YOUR_OLLAMA_API_KEY",
      "contextLength": 16384,
      "completionOptions": {
        "temperature": 0.1,
        "maxTokens": 2048
      }
    },
    {
      "title": "DeepSeek Coder 7B (VPS)",
      "provider": "ollama",
      "model": "deepseek-coder-v2:7b",
      "apiBase": "https://ollama.yourdomain.com",
      "apiKey": "YOUR_OLLAMA_API_KEY",
      "contextLength": 16384
    }
  ],

  "tabAutocompleteModel": {
    "title": "Codestral (Autocomplete)",
    "provider": "ollama",
    "model": "codestral:7b",
    "apiBase": "https://ollama.yourdomain.com",
    "apiKey": "YOUR_OLLAMA_API_KEY"
  },

  "embeddingsProvider": {
    "provider": "ollama",
    "model": "nomic-embed-text",
    "apiBase": "https://ollama.yourdomain.com",
    "apiKey": "YOUR_OLLAMA_API_KEY"
  },

  "contextProviders": [
    { "name": "code" },
    { "name": "docs" },
    { "name": "diff" },
    { "name": "terminal" },
    { "name": "problems" },
    { "name": "folder" },
    { "name": "codebase" }
  ],

  "slashCommands": [
    { "name": "share", "description": "Export conversation" },
    { "name": "cmd", "description": "Generate terminal command" },
    { "name": "commit", "description": "Generate commit message" }
  ]
}

Step 5: Using Continue.dev

<code"># Tab completion (inline code suggestions):
# Just start typing — Continue suggests completions
# Press Tab to accept, Esc to dismiss
# Works like Copilot but with your VPS model

# Chat (Ctrl+L or click Continue sidebar):
# Ask: "Explain this function"
# Highlight code → Ctrl+L → automatically adds to context
# Ask: "Refactor this to use async/await"
# Ask: "Write tests for this function"
# Ask: "What's wrong with this code?"

# Add files to context:
# @file:src/database.py — adds file to context
# @codebase — indexes entire codebase for context

# Slash commands:
# /edit — edit selected code with instructions
# /commit — generate commit message from git diff
# /cmd — generate terminal command

# Keyboard shortcuts:
# Ctrl+L   — open chat
# Ctrl+I   — inline edit
# Ctrl+Shift+L — add to chat

Step 6: Codebase Indexing (RAG for Code)

<code"># Continue.dev indexes your codebase using embeddings
# Allows asking questions about your entire codebase

# Configure embeddings (already done in Step 4)
# Trigger indexing: Continue sidebar → click gear → "Rebuild Index"

# After indexing, ask codebase-aware questions:
# "@codebase How is authentication implemented?"
# "@codebase Where are database migrations?"
# "@codebase Which files handle API routing?"

# Continue uses the VPS-hosted nomic-embed-text model to embed code chunks
# Searches locally using vector similarity — code never leaves your machine
# (only embeddings are computed on the VPS)

Step 7: Team Configuration (Shared Config)

<code"># Share Continue.dev config across team via version control:
# Add .continue/config.json to your repository (with placeholder API keys)

# For team use, store API key as VS Code setting:
# VS Code Settings → search "continue" → set API key there
# Team config references ${VS_CODE_CONTINUE_API_KEY}

# Or use a team proxy that adds auth:
# Route all Continue traffic through your VPS with auth handled server-side
# Team members use an internal URL without needing personal API keys

Model Performance for Coding

  • Codestral 7B: Best balance for tab completion — fast, accurate for common patterns
  • DeepSeek Coder V2 7B: Strong code generation, good at complex algorithmic problems
  • Qwen2.5 Coder 7B: Excellent multilingual coding, particularly strong for Chinese documentation
  • Codestral 22B: Near-Copilot quality, requires 16 GB RAM VPS

Getting Started

Continue.dev with Ollama on a 8 GB Ubuntu VPS at VPS.DO running Codestral 7B provides a capable coding assistant at $0/month in model API fees. Tab completions generate in 1–3 seconds on CPU-only inference — slower than Copilot but acceptable for deliberate completion acceptance. For interactive chat (code explanation, refactoring), the latency is less noticeable. Upgrade to a 16 GB VPS for Codestral 22B for near-Copilot quality.

Conclusion

Continue.dev with VPS-hosted Ollama delivers a private AI coding assistant — tab completion and chat — where your code never leaves your infrastructure. The configuration takes 15 minutes: install Ollama, pull a code model, expose it via Nginx with auth, and configure Continue.dev’s config.json. For teams handling proprietary code or in regulated industries where sending code to GitHub/OpenAI is prohibited, self-hosted Continue.dev is the practical alternative to GitHub Copilot.

Fast • Reliable • Affordable VPS - DO It Now!

Get top VPS hosting with VPS.DO’s fast, low-cost plans. Try risk-free with our 7-day no-questions-asked refund and start today!