GitHub Copilot CLI With Ollama: Local Models and the New Sandbox

TL;DR
Copilot CLI 1.0.94 finds local Ollama models from /model and ships an OS-level sandbox. Setup, the 4k context trap, and why local is not offline.
To run GitHub Copilot CLI on Ollama, update to CLI 1.0.94 or later, start Ollama with at least a 64k context window, then pick the model from /model. Turn on /sandbox enable too: a local model changes where inference runs, not what the agent's shell commands can touch.
Official Sources#
| Resource | What it confirms |
|---|---|
| Discover local models in GitHub Copilot CLI (GitHub Changelog, Oct 7, 2026) | /model discovery of a running Ollama instance from CLI 1.0.94-0, tool calling and streaming required, local model does not mean offline |
| Local sandboxing for GitHub Copilot now generally available (GitHub Changelog, Oct 7, 2026) | Sandbox GA in Copilot CLI, the Copilot app and VS Code Agent Host, powered by MXC, no extra cost |
| Adding LLM models to GitHub Copilot CLI (GitHub Docs) | COPILOT_PROVIDER_* variables, model requirements, COPILOT_OFFLINE=true |
| About cloud and local sandboxes (GitHub Docs) | Per-OS backends and requirements, what the sandbox does not cover, cloud sandbox meters |
| Copilot CLI integration and Context length (Ollama Docs) | ollama launch copilot, recommended models, default context by VRAM |
| github/copilot-cli releases | 1.0.94 is the current stable release (Oct 8, 2026) |
What shipped on October 7#
GitHub's October 7 event put three related things in front of terminal users at once:
- Local model discovery in Copilot CLI. From version 1.0.94-0,
/modellists models from a running local Ollama instance next to your configured models and GitHub's hosted ones. Discovery does not add anything on its own: you pick a model, review its provider and endpoint, then choose Add and use for this session or Add without switching. No restart needed. It does not install Ollama or pull models, and the model must support tool calling and streaming (changelog). - Local sandboxing, generally available. Commands Copilot runs execute with restricted access to the filesystem, network and credentials, enforced by Microsoft eXecution Container (MXC). It is included with Copilot at no additional cost (changelog).
- Local and cloud routing, announced but not shipped. Microsoft's Command Line post says that "by the end of the month" Copilot's Auto mode will decide per task whether to use on-device or cloud inference, and introduces a local build of MAI Code 1.1 Flash for Windows.
The before state matters here. Until this release, a custom provider in the CLI was environment variables only, and in-session /model showed exactly the one model you configured. That limitation is still the subject of an open request, issue #4358. If you read our earlier Copilot CLI, BYOK and AI credits breakdown, this is the piece that makes the local half of BYOK usable day to day.
Should you run Copilot CLI on a local model?#
| If you need | Use | Why |
|---|---|---|
| Code that never leaves the machine | Local Ollama model plus COPILOT_OFFLINE=true | Offline mode stops the CLI calling GitHub; a local provider keeps prompts on the box |
| The strongest agent on hard multi-file work | GitHub's hosted models | A laptop-sized local model trails frontier models on long agent loops |
| A cheap second opinion or bulk chores (renames, test scaffolds, docs) | Local model through /model, switched per session | No per-token bill, and switching back to a hosted model needs no restart |
| Safer autonomy regardless of model | /sandbox enable | Sandbox policy applies to tool execution whichever model is driving |
| One harness across many local runtimes (LM Studio, vLLM, Foundry Local) | The COPILOT_PROVIDER_BASE_URL route | /model discovery is Ollama only; other OpenAI-compatible servers still go through env vars |
Picking the runtime itself is a separate decision. Our Ollama vs LM Studio vs vLLM vs llama.cpp comparison covers that, and the best local coding LLMs roundup covers which weights are worth the disk space.
Prerequisites#
- Copilot CLI 1.0.94 or later. Check with
copilot --version. On our machine the freshly installed package printedGitHub Copilot CLI 1.0.94. - Ollama installed and running, with a model that supports tool calling and streaming already pulled.
- Memory for the context window. Ollama defaults to 4k context under 24 GiB of VRAM, 32k from 24 to 48 GiB, and 256k at 48 GiB or more (Ollama docs). That default is the single most common reason a local Copilot session goes wrong (see Troubleshooting).
- For the sandbox on Linux:
bwrap0.5.0 or later on yourPATH, plusslirp4netns, util-linux 2.35+ and theiptablestooling. macOS uses Seatbelt (macOS 15 or later is the tested floor). Windows needs a recent Windows 11 build (GitHub docs). - A Copilot account, probably. GitHub's install page lists an active Copilot subscription as a prerequisite and says the CLI is available on all Copilot plans. The CLI's own
copilot help providerstext, however, says "GitHub authentication is not required when using a custom provider." The two do not agree; plan on having an account, and treat sign-in-free BYOK as something to confirm for your setup.
Steps#
- Install or update the CLI. Any of GitHub's documented routes works:
npm install -g @github/copilot # Node.js 22+
brew install --cask copilot-cli # macOS and Linux
winget install GitHub.Copilot # Windows
copilot --version # want 1.0.94 or later- Give Ollama a context window an agent can use. Ollama recommends at least 64k tokens for Copilot. Set it when you start the server, then confirm what was actually allocated:
OLLAMA_CONTEXT_LENGTH=64000 ollama serve
ollama ps # the CONTEXT column should read 64000 or more, PROCESSOR ideally 100% GPU
- Pull a model that can call tools. Ollama's Copilot page recommends
qwen3.5andglm-4.7-flashas local models. Everything else on that list ends in:cloud, which runs on Ollama's hosted service, not your machine.
ollama pull qwen3.5
-
Pick it from
/model. Startcopilot, run/model, choose the discovered Ollama model, check the provider and endpoint, then choose Add and use for this session. If the connection fails, the picker shows the reason inline. -
Or wire it up explicitly. On builds before 1.0.94, for scripts, or for a non-Ollama server, use the BYOK variables. This is the Ollama example from
copilot help providersin 1.0.94:
COPILOT_PROVIDER_BASE_URL=http://localhost:11434/v1 \
COPILOT_MODEL=qwen3.5 \
copilot
Ollama also ships a one-line launcher, ollama launch copilot, and documents a headless form for CI: ollama launch copilot --model <model> --yes -- -p "how does this repository work?". The --yes flag pulls the model and skips selectors.
- Go fully offline if that is the point.
COPILOT_OFFLINE=truemakes the CLI skip GitHub authentication, telemetry, the GitHub MCP server, auto-update and its network-reaching tools, and it requires a custom provider (percopilot help environment). It only isolates you if the provider is local too.
COPILOT_OFFLINE=true \
COPILOT_PROVIDER_BASE_URL=http://localhost:11434/v1 \
COPILOT_MODEL=qwen3.5 \
copilot
- Turn on the sandbox. Inside a session,
/sandbox enablepersists across future sessions until/sandbox disable. For a single run, start withcopilot --sandbox. Shell commands, file search and, by default, local MCP and language servers then run inside the OS sandbox.
Verify It Works#
copilot --versionprintsGitHub Copilot CLI 1.0.94.or later.ollama psshows your model loaded with the context you set.- In the session,
/modelshows the Ollama model as the active one. Ask for something that needs a tool, such as "list the files in this repo and summarize the largest one." A model without working tool calls fails here, not on plain chat. - Run
/sandboxto see the effective policy, then ask the agent to write a file outside the working directory. With sandboxing on, the command should fail instead of touching your system.
Troubleshooting: where it breaks#
The agent forgets the task after two turns, or tool calls come back malformed. Check ollama ps. On a GPU with less than 24 GiB of VRAM the default context is 4k. Copilot CLI's fixed overhead is far larger than that: one user's wire capture in issue #2627 measured roughly 16,800 tokens per turn before any repo context, about 9,100 in the system message and 7,300 in tool schemas. The same thread reports CPU-only Ollama setups (tested with a 3B coder model) taking more than five minutes per turn. Set OLLAMA_CONTEXT_LENGTH to 64000 or more, and accept that small CPU-only setups will be slow regardless.
"I picked a local model, so nothing leaves my laptop." Not quite. GitHub is explicit that choosing a local model does not turn on offline mode or disable telemetry, and an Ollama :cloud model sends prompts to Ollama's hosted service. If privacy is the requirement, use a local tag and set COPILOT_OFFLINE=true.
400 errors with a cloud-routed model. Issue #3839, still open, reports Ollama Cloud rejecting the custom_tool_call input items Copilot CLI sends in Fleet Mode, with unknown input item type: "custom_tool_call". That is a compatibility-layer gap, not a model problem.
Your other local server does not appear in /model. Discovery covers a running Ollama instance only. LM Studio, vLLM and Foundry Local still use COPILOT_PROVIDER_BASE_URL, and with that route you get the one configured model.
The sandbox is on, but a secret still leaked into a command. The CLI's help sandbox text notes that sandboxed commands inherit your shell environment apart from a fixed blocklist, so something like AWS_ACCESS_KEY_ID stays visible unless you configure it for masking. The CLI's built-in file tools also run in-process and honor the policy "on a best-effort basis" rather than under OS enforcement. GitHub describes the whole thing as lighter-weight isolation, not a VM or container.
/sandbox enable is refused on Linux. Install or upgrade bubblewrap. The host probe only checks for bwrap, so if a sandboxed command then fails on startup, the missing piece is usually slirp4netns, /dev/net/tun access, or the nf_tables iptables backend.
What people are actually saying#
The demand for this was visible in the issue tracker long before October 7. Issue #4358 asked for /model to read the provider's /models endpoint so BYOK users could switch without restarting, and a second team chimed in to say it blocks them too. Discovery answers that for Ollama only, and the issue is still open for everyone else.
The skeptical case is about overhead, not features. The author of issue #2627 argues that a fixed system prompt and tool schema in the tens of thousands of tokens makes local providers impractical, and asks for a slimmer, configurable prompt. Nobody from GitHub has replied there yet.
The counter-case comes from Microsoft's own numbers, which are vendor-reported. On a Surface Laptop Ultra with up to 128 GB of unified memory, the quantized local MAI Code 1.1 Flash (137B total, 6.8B active parameters, 53 GB quantized) scored 70.8% on SWE-bench Verified against 72.6% for the full model, with peak memory of 75.5 GB at 256k context. In other words: local coding agents work, if the box has the memory to hold the context the harness needs. That is a very different machine from the median developer laptop.
Who wins, and the second-order effect#
Ollama wins the default. It is the only runtime /model discovers, and Ollama's own docs already carry a copilot launcher. Being the zero-config local path inside GitHub's own terminal agent is distribution most runtimes cannot buy.
Microsoft wins the hardware story. MAI Code 1.1 Flash through the Windows ML provider, Copilot routing to it automatically, and MXC on all three operating systems give Windows a coherent local-agent pitch. Our MAI models explainer has the background on that model family.
The losers are assumptions. "Local means private" was always shaky, and Auto routing makes it shakier: once Copilot decides per task whether to run on-device or in the cloud, you can only reason about data flow through settings like offline mode, not through the model name.
The second-order effect is that the sandbox, not the model picker, becomes the thing teams govern. GitHub says it plainly: "model execution and tool isolation are separate concerns." Enterprise-managed settings can require sandboxing, block weakening it, and with sandbox.failIfUnavailable refuse to run model requests at all on a host that cannot enforce it. Claude Code already sandboxes Bash with Seatbelt and bubblewrap, the same OS primitives MXC uses on macOS and Linux, so OS-level isolation is now the baseline for terminal agents rather than a differentiator. Our agent security models comparison shows how quickly that baseline moved, and the MXC developer guide shows how to use the same library in your own agent.
FAQ#
Can GitHub Copilot CLI use Ollama?#
Yes. From CLI 1.0.94-0, /model discovers models from a running Ollama instance and lets you add one for the current session without restarting. Older builds, and other OpenAI-compatible servers, use COPILOT_PROVIDER_BASE_URL=http://localhost:11434/v1 plus COPILOT_MODEL.
Does using a local model make Copilot CLI offline?#
No. GitHub says choosing a local model does not turn on offline mode or disable telemetry. Set COPILOT_OFFLINE=true for that, and keep the provider local, because a remote provider still receives prompts and code context.
Which Ollama model should I use with Copilot CLI?#
It must support tool calling and streaming. Ollama recommends qwen3.5 and glm-4.7-flash as local options; its :cloud tags run on Ollama's hosted service. Give it at least 64k of context, and GitHub suggests 128k or more for best results.
Is local sandboxing in Copilot free?#
Yes. Local sandboxing is included with GitHub Copilot at no additional cost. Cloud sandboxes are separate, in public preview, and billed per use: $0.000024 per compute second, $0.000003 per GiB-second of memory, and $0.005 per GiB-month of snapshot storage.
Does the Copilot sandbox work on Linux?#
Yes, through bubblewrap. You need bwrap 0.5.0 or later on your PATH, plus slirp4netns, util-linux 2.35+, the iptables tooling and access to /dev/net/tun. Enable it with /sandbox enable, or copilot --sandbox for one session.
Sources#
Checked October 9, 2026.
Last updated: October 9, 2026
Continue Reading#
- GitHub Copilot CLI, BYOK, and AI Credits: The New Cost-Control Stack - the June release that introduced bring-your-own-key, and what it means for cost.
- Ollama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents - when Ollama is the right local server and when it is not.
- Microsoft MXC Developer Guide 2026: Sandbox Your AI Agents at the OS Level - the library under Copilot's sandbox, usable in your own agents.
- AI Coding Agent Security Models Compared 2026 - permissions, sandboxing and credential handling across every major coding agent.
- The Best Local Coding LLMs in 2026 - which open weights hold up on agent work on your own hardware.
Get the next deep dive like this in your inbox
One email a week on GitHub Copilot and the rest of the AI dev stack. Free.
Read next on local and open-weight models
GitHub Copilot CLI, BYOK, and AI Credits: The New Cost-Control Stack
GitHub's June Copilot updates point beyond autocomplete: CLI access, bring-your-own-key model routing, AI credit metrics, and external agent providers make Copilot a governed agent platform.
8 min readMicrosoft MXC Developer Guide 2026: Sandbox Your AI Agents at the OS Level
Microsoft Execution Containers (MXC) give your AI agents policy-driven sandboxing across Windows, Linux, and macOS. TypeScript SDK, JSON config, multiple isolation backends. Here is how to use it.
8 min readOllama vs LM Studio vs vLLM vs llama.cpp: Picking a Local Runtime for Coding Agents
A fair, sourced comparison of the four runtimes developers reach for when they want a coding agent talking to a model on their own hardware instead of an API: Ollama's convenience, LM Studio's GUI, vLLM's throughput, and llama.cpp's control. What each is actually for, and which to pick.
10 min readNew here? Start with
Technical content at the intersection of AI and development. Building with AI agents, Claude Code, and modern dev tools - then showing you exactly how it works.







