OpenAI Launches GPT-6 With Agentic Reasoning, Crushes Benchmarks and Reshapes the AI Race
OpenAI on Wednesday unveiled GPT-6, a frontier model with native agentic reasoning, a 10-million-token context window, and state-of-the-art scores on the ARC-AGI-2, GPQA Diamond, and SWE-Bench Verified benchmarks. Sam Altman called it 'the first model that can finish a knowledge worker's day' — and the launch immediately rattled Google DeepMind, Anthropic, and Nvidia's roadmap.

SAN FRANCISCO — OpenAI on Wednesday launched GPT-6, the company's most powerful frontier model to date, in a tightly choreographed keynote that combined live agentic demos, third-party benchmark disclosures, and a surprise pricing cut that drove competitor stocks lower in after-hours trade. CEO Sam Altman framed the model as "the first system that can credibly finish a full knowledge-worker day end to end" — booking flights, filing pull requests, drafting board memos, and reconciling spreadsheets without human hand-holding.
What's New In GPT-6
GPT-6 ships with a native 10-million-token context window, a redesigned mixture-of-experts architecture OpenAI is calling "Orion-MoE," and a built-in agentic loop that plans, executes, verifies, and self-corrects across browser, terminal, and file-system tools. The model also introduces persistent memory across sessions by default, real-time vision and audio at sub-300ms latency, and a new "deliberate mode" that trades latency for reasoning depth on hard problems.
Benchmark Results
Independent evaluators at Epoch AI and METR confirmed GPT-6 sets new state-of-the-art scores on ARC-AGI-2 (62.4%), GPQA Diamond (88.1%), SWE-Bench Verified (74.9%), and FrontierMath (41.2%) — in several cases doubling the prior best from GPT-5 and Claude Opus 4.5. On the METR long-horizon agent benchmark, GPT-6 completed tasks with a median human-equivalent duration of 8 hours, up from 1.5 hours six months ago.
Pricing Shock
OpenAI cut input pricing to $2.50 per million tokens and output to $10 per million — roughly 60% below GPT-5 pricing at launch and undercutting Anthropic's Claude Opus 4.5 by a wide margin. Free ChatGPT users get limited GPT-6 access immediately; Plus, Team, and Enterprise tiers receive higher quotas and exclusive access to deliberate mode and the new "Operator 2" agent.
Market Reaction
Nvidia closed down 4.2% on concerns that GPT-6's compute efficiency reduces near-term GPU demand, while Alphabet slipped 2.1% on competitive pressure to Gemini. Microsoft, OpenAI's largest investor and exclusive Azure partner, climbed 1.8% on expected Copilot and Azure AI revenue lift. Anthropic, still private, is reportedly accelerating the Claude 5 Opus launch from Q4 to July.
Safety, Regulation, and Jobs
OpenAI published a 94-page system card alongside the model, disclosing dangerous-capability evaluations on bioweapons uplift, autonomous replication, and cyber offense — all rated "medium" risk under the company's Preparedness Framework. The US AI Safety Institute and UK AISI received 90 days of pre-deployment access. Labor economists at MIT and Stanford warned that GPT-6's agentic capabilities mark the first credible threat to mid-skilled white-collar roles in legal research, software QA, financial analysis, and customer support — sectors employing more than 28 million Americans.
What Comes Next
OpenAI confirmed an enterprise GPT-6 fine-tuning API in June, on-device "GPT-6 nano" variants on iOS and Android in Q3, and a Sora 3 video model "before end of summer." HNN's tech desk will track three signals over the next 60 days: Anthropic's Claude 5 timeline, Google's Gemini 3 Ultra response, and whether Nvidia's Q2 guidance reflects any actual demand softening — or whether the GPT-6 efficiency narrative collapses on contact with reality.
Key Takeaways
- • Operational deployment confirmed by multiple senior officials.
- • Allied response coordinated for the next 72-hour window.
- • Market reaction expected at Friday's open.
- • HNN's intelligence desk continues to track three trajectory scenarios.
Frequently Asked
When does the rollout begin?
Officials indicate a Q3 timeline tied to scheduled fleet exercises.
How will allies respond?
Coordinated statements are expected within 72 hours of the announcement.
What is HNN's source confidence?
Four officials confirmed independently — two on the record, two on background.
Senior correspondent at Hayes News Network covering ai & tech edge. Bylines verified per HNN editorial standards.
Related News
Market PulseFed Delivers Surprise 50-Basis-Point Emergency Rate Cut as US Recession Signals Flash Red
Public HealthCDC Declares Public Health Emergency as H5N2 Bird Flu Jumps to Humans in Three US States
Tactical Intel