Plan
Use GPT-6.1 Sol or another stronger model to clarify goals, compare approaches, and write testable steps.
Deliverable: a bounded plan + acceptance checks.A compact reference for choosing models in Copilot Chat and agent workflows. Use it to balance model quality, token cost, latency, and the type of work you are asking Copilot to do.
A starting strategy, not a rule: spend more where judgment matters and less where the steps are clear.
Use GPT-6.1 Sol or another stronger model to clarify goals, compare approaches, and write testable steps.
Deliverable: a bounded plan + acceptance checks.Try GPT-6 Luna for focused edits, documentation, and implementation against that plan.
Deliverable: small changes + passing checks.Use a stronger model for risky decisions or unresolved failures. Verify results with tests and human judgment.
Escalate when the same failure repeats.Give it the goal, relevant files, the agreed approach, constraints, and commands or checks that prove success. Work in small steps. Ask it to stop and report a blocker instead of guessing.
“Implement step 2 of this plan. Keep the public API unchanged. Run the listed checks. If the assumptions fail, explain what you found before changing the approach.”
Switch models in the Copilot model picker. Preserve the plan and relevant context when starting a new conversation; switching alone does not guarantee an effective handoff.
Choose a situation to see a practical starting point.
These are editorial recommendations, not model benchmarks. Judge the actual result, not the model’s confidence.
20,000 uncached input tokens + 4,000 output tokens, at standard rates.
Linear scale, $0 to $0.400. Gemini 3.8 Flash uses its promotional rate through December 31, 2026. This compares prices for equal tokens, not quality or cost per successful task. Input tokens are what you send; output tokens are what the model generates, including billed reasoning.
GitHub’s official rates · AI-credit billing: 1 credit = $0.01. Subscription fees, included allowances, caching, and long-context tiers are excluded here. Availability depends on your plan, client, and organization policy. Legacy annual request-based plans differ.
Adjust the work. See what model switching could save.
Illustrative estimate, not a benchmark. Both routes include the same planning + review steps. Every step uses the token counts you enter; real steps vary.
Cost = (input tokens × input rate + output tokens × output rate) ÷ 1,000,000. No cache reads or writes. Standard tier only; per-step input is capped at 272K. Dollars show usage value, not necessarily an extra charge on your bill.
Start with Luna on a bounded task. If your client offers thinking effort, try a higher setting when deeper reasoning is useful. There is no published universal conversion from “High” to a fixed cost.
Check the result. If it repeats a failed approach, misses constraints, or needs architectural judgment, pause and switch. A cheap attempt that never finishes is not an efficient workflow.
Read more: GitHub’s model comparison · Full pricing reference
Enterprise Routing Graphic
A current October 10, 2026 view of five selected Copilot Enterprise routes plus GPT-5 mini migration guidance. Prices are GitHub Copilot AI-credit rates per 1M tokens; 1 AI credit equals $0.01.
Do not start new workflows on GPT-5 mini. Migrate fixed collection and routine prompts to GPT-5.6 Luna before October 19, 2026.
Use for quick checks and lightweight coding loops; it is the lowest-cost GPT-6 route. Keep GPT-5 mini migrations on GitHub's listed GPT-5.6 Luna target until separately validated.
Use for balanced everyday interactive and agentic coding. Keep tests, source checks, and human review as independent validation gates.
Use for well-scoped everyday development such as feature work and bug fixes; GitHub's early testing matched Sonnet 5 coding quality with fewer steps, tokens, and tool calls plus faster completion.
Use for agentic coding and terminal workflows that benefit from strong multistep performance and efficient token use.
Reserve for long-running, high-consequence agent work that needs efficient multistep recovery after cheaper routes fail.
Reference Table
Copilot Enterprise seats contribute a pooled monthly AI-credit allowance. Usage is token-based by model and token type, one AI credit equals $0.01, and usage beyond the included pool is billed in AI Credits at GitHub's listed per-token rates.
| Model | Provider | Status | Input / 1M | Cached / Write | Output / 1M | Use It For |
|---|---|---|---|---|---|---|
| GPT-5 mini | OpenAI | GA · deprecation scheduled Oct 19 | $0.25 | $0.025 cached; no cache write | $2.00 | Currently GA; scheduled for deprecation on October 19, 2026. Migrate routine lightweight workflows to GPT-5.6 Luna. |
| GPT-5.6 Luna | OpenAI | GA | $0.20 | $0.02 cached; $0.25 write | $1.20 | Cheap iteration, quick checks, and lightweight coding loops. |
| GPT-6 Luna | OpenAI | GA | $0.10 | $0.01 cached; $0.125 write | $0.50 | Lowest-cost GPT-6 route for smaller, faster tasks; gradual rollout and client or policy availability can vary. |
| GPT-5.3-Codex | OpenAI | GA · base + LTS | $1.75 | $0.175 cached; no cache write | $14.00 | Agentic software development; the Business and Enterprise base model when no other model is enabled, with LTS availability through February 4, 2027. |
| GPT-5.6 Terra | OpenAI | GA | $2.00 | $0.20 cached; $2.50 write | $12.00 | Balanced everyday interactive and agentic coding; use tests, source checks, and human review for validation. |
| GPT-5.6 Sol | OpenAI | GA | $4.00 | $0.40 cached; $5.00 write | $20.00 | Long multipart orchestration and diagnosis when the higher full-rate cost is justified. |
| GPT-6.1 Sol | OpenAI | GA | $2.00 | $0.10 cached; $2.50 write | $10.00 | Agentic coding and terminal workflows with strong multistep performance and efficient token use; gradual rollout and client or policy availability can vary. |
| GPT-6 Astra | OpenAI | GA | $10.00 | $1.00 cached; $12.50 write | $50.00 | Premium OpenAI route for long-horizon autonomous coding when planning and validation quality outweigh cost. |
| Claude Haiku 5.5 | Anthropic | GA | $0.10 | $0.01 cached; $0.125 write | $0.50 | Fast, high-volume work such as subagents, quick edits, and terminal tasks; gradual rollout and plan or policy availability can vary. |
| Claude Sonnet 5.5 | Anthropic | GA | $2.00 | $0.10 cached; $2.50 write | $10.00 | Well-scoped everyday feature work and bug fixes with efficient task completion; gradual rollout and plan or policy availability can vary. |
| Claude Opus 5 | Anthropic | GA | $5.00 | $0.50 cached; $6.25 write | $25.00 | Exceptional unresolved or high-consequence escalation after cheaper models fail. |
| Claude Opus 5.5 | Anthropic | GA | $4.00 | $0.20 cached; $5.00 write | $20.00 | Long-running agentic coding and knowledge work with efficient multistep recovery; gradual rollout and plan or policy availability can vary. |
| Claude Fable 5.1 | Anthropic | GA | $10.00 | $0.25 cached; $12.50 write | $50.00 | Specialized premium route only after retention review and administrator enablement; Enterprise ZDR requires approved access. |
| Gemini 3.8 Flash | GA promo | $0.75 promo | $0.075 cached; no cache write | $3.75 promo | Cost-aware Google route for complex terminal coding tasks during introductory pricing through December 31, 2026. |
Enterprise billing caveat: Auto can route around busy or degraded models, and Copilot code review automatically selects an undisclosed model. Models also vary by client, enterprise policy settings, preview access, context size, and reasoning level. Rates shown are default-tier rates; long-context usage can cost more. Legacy annual-plan multiplier language for some individual plans is not the Enterprise billing model; Selected-model usage is metered by input, cached input, cache write where applicable, and output tokens; Copilot code review also consumes GitHub Actions minutes.
Auto-selection option: GitHub offers Efficiency, Balance, and Intelligence tiers in VS Code, Copilot CLI, and the Copilot app. Efficiency favors low cost, Balance weighs cost, quality, and speed, and Intelligence favors quality for complex work. All tiers use the same eligible model pool; Auto evaluates each prompt, so even Intelligence can select a small model for a simple task. Usage follows the selected model's token rate regardless of tier, with a 10% Auto discount for paid plans. Use Auto when policy allows it, but select a model explicitly when retention, capability, or predictable per-token cost matters.
October 2 migration: GitHub deprecated Claude Opus 4.7, Gemini 3.5 Flash, Gemini 3.6 Flash, and Kimi K2.7 Code across Copilot on October 2, 2026. Move explicit selections to Claude Opus 5.5, Gemini 3.8 Flash, or Kimi K3 as applicable, and verify current availability in the model picker and organization policy.
October 19 migration: GitHub has scheduled GPT-5 mini, GPT-5.4 mini, GPT-5.4, GPT-5.5, Gemini 3.7 Flash, and Grok 4.5 for deprecation across Copilot. Move explicit selections to GitHub's listed replacements before the deadline and confirm those models are enabled by Enterprise policy.
Fable governance caveat: Claude Fable 5 and 5.1 retain prompts and outputs by default and are excluded from default model enablement. Enterprise zero-data-retention access requires GitHub approval and configuration under the current time-limited exception through the end of 2026.
HydraFusion research-preview caveat: Copilot CLI, VS Code 1.140 or later, and the GitHub Copilot app can use HydraFusion to choose among single, cascade, and independent-critique patterns from a model pool spanning multiple providers. Business and Enterprise administrators must allow preview features. Usage is billed from every model leg at its standard token rate, so treat benchmark savings as directional, inspect actual cost and quality on your own tasks, and keep normal validation and policy gates.
Task Guide
These are site recommendations grounded in GitHub's published capabilities and pricing and OpenAI's model-selection guidance. This site takes a conservative approach for new or high-risk workflows: establish an accuracy baseline with the most capable model, then compare identical inputs and test smaller models for cost and latency. For established work, keep the lightest model and settings that consistently meet the quality bar.
Start: GPT-6 Luna for fixed Python/API collection, known-schema edits, and cheap iteration.
Escalate: GPT-5.6 Terra when the edit needs balanced interactive coding, with independent tests and source checks.
Start: GPT-5.6 Terra for balanced everyday interactive and agentic coding.
Escalate: GPT-6.1 Sol for multipart orchestration and efficient multistep validation.
Start: GPT-5.6 Terra for a balanced debugging pass, then verify logs, repro evidence, and claims independently.
Escalate: GPT-6.1 Sol for long diagnosis, then Claude Opus 5.5 only if the issue remains unresolved.
Start: GPT-6.1 Sol for long multipart planning and cross-file coordination when Terra is not enough.
Escalate: Claude Opus 5.5 for exceptional high-consequence review after cheaper routes fail.
Start: GPT-5.6 Terra for a focused everyday review, backed by tests, source checks, and human judgment.
Escalate: Claude Opus 5.5 for high-risk logic, security-sensitive areas, or architecture drift.
Start: GPT-6 Luna for source collection and first drafts from known evidence.
Escalate: Claude Sonnet 5.5 only when blind review shows it improves audience-fit communication.
Start: Confirm visual input is supported in the active Copilot client and choose a currently supported multimodal model; do not build a new visual workflow on GPT-5 mini because it is scheduled for deprecation on October 19, 2026.
Escalate: Convert the visual evidence into explicit text requirements before moving to a deeper model whose active client supports the needed inputs.
Start: GPT-6 Luna before using Terra, Sol, Sonnet 5, or Opus 5.5.
Escalate: Spend higher token-rate models only after the question, evidence, and acceptance criteria are narrowed.
Recommended Flows
Use different models for different phases instead of trying to make one model do every job.
Cost Notes
Copilot Enterprise usage is token-based. The real cost depends on prompt length, cached context, cache writes where applicable, outputs, selected model, and agent surface.
GPT-6 Luna is the lowest-cost current GPT-6 route for routine prompts. GPT-5 mini remains available today but is scheduled for deprecation on October 19, 2026; follow GitHub's listed migration target of GPT-5.6 Luna for existing workflows, then evaluate GPT-6 Luna separately before changing production behavior. GPT-4.1 is retired from selectable Copilot use, though GitHub still lists it as a background utility model; move workflows that selected it explicitly to a supported model.
GPT-5.6 Terra, GPT-6.1 Sol, and Claude Sonnet 5.5 sit in the middle of this six-model route, but they serve different jobs: Terra for balanced everyday interactive and agentic coding, GPT-6.1 Sol for efficient complex agentic coding, and Sonnet 5.5 when blind tests favor its everyday development, CLI, latency, or communication fit.
GPT-6 Astra and Claude Opus 5.5 should answer narrow, high-value questions rather than carrying every iteration of a long coding session. Consider Claude Fable 5.1 only after its retention and ZDR requirements are approved.
Sources
Verified October 10, 2026. The model list, context-window controls, reasoning controls, agent-app surfaces, sandbox policy controls, security-validation defaults, model data-retention requirements, managed-settings coverage, default model enablement, enterprise team targeting, VS Code model-provider controls, code-review billing behavior, retirements, and AI-credit billing details change often, so re-check these sources before making budget or policy decisions.