GPT 5.6 Luna on VM0. The fast, economy GPT-5.6 tier
The economy tier of OpenAI's GPT-5.6 preview family. Fast, cheap and multimodal — built for high-volume tool-use, bulk work and latency-sensitive replies.
400K tokens · Text / Vision / Code · Prompt cache
GPT 5.6 Luna is the economy tier of OpenAI's GPT-5.6 preview generation — the fast, cheap member of the family built for high-volume tool-use, bulk classification and extraction, and latency-sensitive chat where speed and cost matter more than frontier reasoning depth.
Vendor list price is $1 / $6 per 1M tokens with cached input at $0.10 / 1M — the lowest in the GPT-5.6 family. It sits in VM0's economy credit band, which makes it the tier to reach for on the wide base of routine work while GPT 5.6 Terra handles the everyday agent and Sol takes only the hardest steps. GPT-5.6 is currently a preview family; GPT-5.5 remains the default OpenAI model until it graduates.
What is GPT 5.6 Luna?
July 2026 preview (economy tier of the GPT-5.6 family) · Economy tier of the GPT-5.6 family, below Sol and Terra. The fast, low-cost option for high-volume and latency-sensitive work.
GPT-5.6 shipped in July 2026 as a three-tier family — Sol, Terra and Luna. Luna is the economy tier: OpenAI positions it for the high-volume, latency-sensitive base of an agent's workload, the role GPT-5.4 Mini played in the previous generation. It keeps the 400K-token context window and reasoning_effort parameter of the family, so it drops into existing Codex agents without changes.
Luna trades top-end reasoning for speed and cost. At $1 / $6 per 1M tokens it is a fraction of Sol's price and the cheapest supported GPT-5.6 tier on VM0. In OpenAI's preview materials it holds up well on straightforward tool-use, classification and extraction, and only falls behind meaningfully on the hard reasoning and long-loop tasks that belong on Terra or Sol anyway.
As with the rest of the preview family the numbers are early and directional. What is stable is Luna's role: it is the tier you run at the wide base of a pipeline — many cheap, fast calls — while promoting the small fraction of genuinely hard steps to Terra and Sol. That layering is what keeps a GPT-5.6 agent both capable and affordable.
What's notable about GPT 5.6 Luna
Headline architecture and capability features.
GPT 5.6 Luna keeps the 400K-token context window, billed at standard input pricing across the entire window. It supports the reasoning_effort parameter, prompt caching where cached input bills at one-tenth the input rate ($0.10 / 1M) plus a cache-write charge ($1.25 / 1M), and the Responses API surface the codex CLI uses by default. Tool-use, structured outputs and computer-use match the rest of the GPT-5.6 family. Inputs are multimodal across text, vision and code; there is no native image generation.
Specs at a glance
GPT 5.6 Luna benchmarks
Preview figures from OpenAI's GPT-5.6 materials, shown against the public GPT-5.4 Mini numbers Luna succeeds. Treat all percentages as directional: the family is in preview and OpenAI has flagged SWE-bench Verified contamination across frontier models.
GPT 5.6 Luna pricing
Provider list price, per 1M tokens.
How GPT 5.6 Luna behaves in practice
Observed behaviour from production agent runs.
Speed
The fastest tier in the GPT-5.6 family — around 130 tokens/sec at medium effort in early testing. This is the reason to pick Luna: interactive replies and high-throughput pipelines stay responsive.
Tool routing
Reliable on straightforward and moderately structured tool calls. It loses ground to Terra on conditional selection and to Sol on tool calls dispatched after long reasoning — keep those steps on the higher tiers.
Bulk classification and extraction
The right tier for running the same structured prompt across thousands of items. Multimodal input means it handles screenshot- and document-driven extraction, not just text.
Reasoning depth
Lighter than Terra and Sol on hard multi-step reasoning and long agentic loops. Use it for breadth, not depth: many cheap fast calls, with the hard steps promoted upward.
Cost profile
$1 / $6 per 1M and the economy credit band make Luna the cheapest GPT-5.6 tier — the model you run at the wide base of a pipeline where volume, not frontier reasoning, drives the bill.
Best agent tasks for GPT 5.6 Luna
High-volume classification and tagging
Run Luna across thousands of tickets, emails or documents to label, route or extract fields. At $1 / $6 the per-item cost stays low, and multimodal input covers screenshot- and PDF-driven jobs.
Latency-sensitive chat replies
For a user-facing assistant where response time is felt directly, Luna's ~130 tokens/sec keeps replies snappy while still handling tool calls and structured outputs.
The base layer under Terra and Sol
In a layered agent, Luna handles the many cheap, fast steps — fetching, formatting, simple tool calls — while Terra runs the everyday logic and Sol takes the hardest planning. The layering keeps the whole pipeline affordable.
Draft-then-refine loops
Let Luna produce fast first drafts — outlines, candidate code, summaries — and promote only the ones that need polishing to Terra or Sol. You pay the frontier rate on a fraction of the volume.
When to skip GPT 5.6 Luna
Skip Luna on the hardest multi-file refactors, long orchestration loops and graduate-level reasoning, where Terra and especially Sol do materially better. When first-attempt patch quality or deep reasoning is the point, the economy tier is a false economy.
GPT 5.6 Luna vs other models
GPT 5.6 Luna vs GPT 5.6 Terra
Terra is the balanced everyday default; Luna is the economy tier below it. Luna is faster and roughly a third of Terra's list price, but gives up reasoning depth and hard tool-routing. Run Luna for volume and latency; step up to Terra the moment a task needs real reasoning.
GPT 5.6 Luna vs GPT-5.4 Mini
Luna is the GPT-5.6 successor to GPT-5.4 Mini in the economy slot, with the newer generation's tool-use behaviour and multimodal input. It lists slightly higher ($1 / $6 vs $0.75 / $4.50) but brings the GPT-5.6 improvements to the cheap, fast tier.
GPT 5.6 Luna vs Claude Sonnet 4.6
Different families, overlapping economy-to-balanced role. Sonnet 4.6 brings the 1M-token context window and Anthropic's ecosystem; Luna brings the Codex framework, faster generation and a lower price. Pick by framework, context length and how much you value speed over reasoning depth.
Bottom line: should you use GPT 5.6 Luna?
GPT 5.6 Luna is the fast, cheap base of the GPT-5.6 family: run it for high-volume and latency-sensitive work, and promote the genuinely hard steps to Terra and Sol.
Frequently asked questions
What is GPT 5.6 Luna's context window?
400,000 tokens, with up to 128K tokens of output per response. The full window bills at standard rates.
When should I use Luna instead of Terra?
When speed and cost matter more than reasoning depth: high-volume classification and extraction, latency-sensitive chat, and the cheap base layer of a layered agent. Step up to Terra the moment a task needs real multi-step reasoning or hard tool-routing.
Is GPT-5.6 generally available on VM0?
It is in preview. All three tiers are selectable, but GPT-5.5 remains the default OpenAI model until the GPT-5.6 family graduates, and preview benchmark numbers are early.
Does GPT 5.6 Luna support prompt caching?
Yes. Cached input bills at $0.10 per 1M tokens with a cache-write charge of $1.25 per 1M — the cheapest caching in the family. Worth enabling on any agent with a stable prefix.
What framework does GPT 5.6 Luna use on VM0?
Codex. VM0 routes GPT-5.6 through the Codex framework's Responses API surface. Claude Code-framework agents are not compatible with GPT-5 models on VM0.
Alternatives
Using GPT 5.6 Luna on VM0
Two ways to access GPT 5.6 Luna on VM0
VM0 supports GPT 5.6 Luna as a Built-in model billed in VM0 credits, and through bring-your-own with a OpenAI API key. The Built-in path uses VM0 Managed routing and the credit multiplier explained below; the bring-your-own path bills you directly with the upstream vendor and skips the VM0 credit conversion entirely.
VM0's recommendation
VM0 positions GPT 5.6 Luna as a cost-saving option rather than a core agent model. Use it to optimise unit cost on non-core work, such as bulk classification, pre-filters, latency-critical short replies, or pinned legacy agents, while keeping Claude Opus 4.7, Claude Opus 4.6, or Claude Sonnet 4.6 on the steps that decide the run.
Credits and the ×0.4 multiplier
Every Built-in model on VM0 is priced as a multiple of Claude Sonnet 4.6, which sits at the ×1 credit baseline. GPT 5.6 Luna bills at ×0.4 credits. The multiplier is what shows up on your VM0 invoice; the vendor list price in the pricing table above is what the upstream provider charges before VM0 converts it into credits.
GPT 5.6 Luna bills at ×0.4, which means a step here costs only 0.4× the credits of an equivalent step on Sonnet 4.6 (the ×1 baseline). That puts it well below the credit baseline and makes it the natural pick for high-volume background work where cost-per-step matters more than peak reasoning quality.
Available on VM0 since July 2026 (preview).