OpenAI on July 30 introduced main API value reductions for 2 fashions in its GPT-5.6 household, reducing the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%.
GPT-5.6 Luna, the corporate’s quickest and lowest-cost mannequin, now prices $0.20 per million enter tokens and $1.20 per million output tokens. Terra, positioned as the center choice for on a regular basis workloads, drops to $2 per million enter tokens and $12 per million output tokens.
OpenAI CEO Sam Altman introduced the transfer on X, calling them “major price cuts today,” together with the 80% discount for Luna and the 20% discount for Terra. The firm mentioned the reductions stem from enhancements throughout its fashions, inference programs, and agent instruments, enabling GPT-5.6 fashions to finish duties extra effectively.
High-volume intelligence at low-cost margins
The pricing shift strikes GPT-5.6 Luna immediately into competitors with lower-cost inference choices from rival suppliers, together with Google’s Gemini 3.5 Flash-Lite and choices from Chinese startups.
According to benchmarking cited by OpenAI, Luna matches the capabilities of frontier-class fashions from a 12 months prior at roughly 6 cents on the greenback per activity whereas executing practically 9 instances quicker. On the skilled benchmark Agents’ Last Exam, OpenAI claims Luna outperforms Anthropic’s Claude Fable 5 “at an estimated cost per task nearly 99% lower.”
OpenAI attributed the worth cuts to inside effectivity good points pushed autonomously by GPT-5.6 Sol itself. Within a human-led engineering course of, Sol modified manufacturing software program kernels and ran token-generation experiments. These automated changes decreased end-to-end mannequin serving prices by 20% and boosted token-generation effectivity by over 15%, based on the corporate.
The shift from mannequin entry to unit economics
As enterprise adoption shifts from experimental deployment to high-volume operational workflows, AI builders are dealing with heightened scrutiny concerning return on funding.
The steep drop in per-token prices displays an evolving market technique the place entry to baseline intelligence is quickly commoditizing.
Rather than competing solely on uncooked efficiency metrics on the prime finish of the portfolio, AI distributors are more and more competing on unit economics to retain enterprise purchasers working steady, multi-step agentic programs.
For company expertise patrons, these changes make complicated agent workflows corresponding to multi-stage code era, doc parsing, and automatic routing economically viable at scale.
By aggressively decreasing the ground on light-weight fashions like Luna whereas charging a premium for latency-focused processing on Sol, OpenAI is encouraging clients to orchestrate mixed-model architectures: assigning lower-cost nodes to deal with routine steps whereas reserving high-tier reasoning for complicated bottlenecks.
Check out our AI pricing information evaluating present prices for ChatGPT, Claude, Gemini, picture mills, video instruments, and voice AI.


