Introduction
OpenAI has officially merged ChatGPT and Codex, and the highly anticipated GPT-5.6 series is now live. The once-central chat box in ChatGPT has been replaced by a Codex-style workspace, marking the end of the pure chat era and the beginning of a new AI age.
The Three Models of GPT-5.6
The GPT-5.6 series includes three models:
- Sol: The new flagship model, designed for complex reasoning, coding, research, and long-horizon agentic work.
- Terra: A balanced model for efficient everyday work, with performance comparable to GPT-5.5 at a lower cost.
- Luna: The fastest and most cost-efficient model, ideal for high-volume tasks.
All three models are available to all ChatGPT subscription plans, making advanced AI capabilities accessible to every member.
GPT-5.6 Sol: The All-Rounder Model
After five hours of intensive use, GPT-5.6 Sol stands out as the most well-rounded model available today. It offers low pricing, strong planning capabilities, exceptional code execution, extremely low hallucination rates, and high accuracy. While its front-end aesthetics and content creation still have room for improvement, it outperforms previous models in these areas.
Key Performance Highlights
- Coding: GPT-5.6 Sol scores 80 on the Artificial Analysis Coding Agent Index, surpassing Claude Fable 5's 77.2 by 2.8 points.
- Efficiency: It reduces output token latency by more than half and lowers costs by one-third compared to other leading models.
- Cybersecurity: GPT-5.6 Sol achieves results comparable to Claude Fable 5 on the ExploitBench cybersecurity evaluation at a fraction of the cost.
Real-World Application: Cybersecurity
For an AI news website under frequent DDoS attacks, GPT-5.6 Sol proved invaluable. When using Claude Fable 5 for security tasks, it was often restricted due to security keywords. GPT-5.6 Sol, however, identified and fixed multiple security vulnerabilities in just 21 minutes using its Ultra Fast mode, with minimal token consumption.
User Experience Improvements
GPT-5.6 Sol also shows significant improvements in front-end aesthetics. For example, visualizing data for Typhoon Bavi produced results far better than what GPT-5.5 could produce, though still not quite at the level of Claude's best outputs.
The ChatGPT and Codex Merge
OpenAI has transformed ChatGPT into ChatGPT Work, making it easier for the 1 billion weekly active ChatGPT users to transition to Codex's capabilities. Users can now switch between Work and Codex modes in the top-left corner, with the original chat mode now located in the sidebar.
Conclusion
GPT-5.6 represents a significant leap forward in AI capabilities, offering powerful, cost-effective solutions for coding, cybersecurity, and everyday work. The merge of ChatGPT and Codex, along with the new model lineup, marks the beginning of the agentic AI era, providing users with more powerful and versatile tools to tackle complex tasks.
常见问题
What does the ChatGPT-Codex merger mean for existing ChatGPT users?
ChatGPT is now called "ChatGPT Work" — the familiar chat interface still exists in the sidebar, but the default view is now a Codex-style workspace. This means every ChatGPT user now has access to agent capabilities (file operations, code execution, browser control) without installing a separate app. The transition is designed to be gradual: you can still use pure chat mode for simple questions, but the workspace is there when you need the AI to actually do things rather than just talk about them. For 1 billion weekly active users, this is the biggest interface change since ChatGPT launched.
How significant is Sol beating Claude Fable 5 on coding benchmarks?
Very. Claude Fable 5 was the undisputed coding champion, and Sol's 80 vs. 77.2 on the Coding Agent Index is a meaningful gap — not just a statistical tie. More importantly, Sol achieves this while reducing latency by 50%+ and cost by one-third. This combination — better quality, faster speed, lower cost — is rare in AI benchmarks. The cybersecurity results are equally notable: matching Claude's security analysis quality at a fraction of the cost makes Sol the practical choice for security-sensitive development work. The one area where Claude still leads is front-end aesthetics and creative output quality.
Should I switch from Claude to GPT-5.6 Sol?
It depends on your primary use case. For coding and agentic workflows, Sol is now the stronger choice — better benchmarks, lower cost, faster output. For front-end design and creative content, Claude still produces more polished results. For security-sensitive work, Sol's ability to handle security tasks without triggering keyword restrictions (unlike Claude) is a practical advantage. The smart approach: use both. Sol for heavy coding, agent tasks, and security work. Claude for design, creative writing, and when you need the most polished final output. The cost difference means you can use Sol for 80% of tasks and Claude for the 20% where it excels.
Is Luna worth using, or is it just a stripped-down model?
Luna is purpose-built for high-volume, low-complexity tasks. It's not "stripped down" — it's optimized for a different workload. For tasks like summarizing 100 customer emails, formatting a dataset, or generating routine reports, Luna is actually the better choice than Sol or Terra because it's faster and cheaper while being perfectly adequate for the task. The companion article (GPT-5.6 Family practical testing) shows Luna generating a 47-page report from 26K records — that's real work, not a toy. The key is task matching: don't use Luna for complex reasoning, don't use Sol for simple data processing.