OpenAI Unveils Custom Inference Chip, Sending Nvidia Shares Tumbling
OpenAI took the wraps off its first in-house inference accelerator on Saturday, claiming a 4x cost-per-token advantage over current Hopper GPUs and confirming Broadcom and TSMC as silicon partners. Nvidia slid 6% in after-hours trading as analysts began rewriting their 2027 AI capex models in real time.

SAN FRANCISCO — OpenAI on Saturday took the wraps off its first in-house inference accelerator, a chip the company says delivers a 4x cost-per-token advantage over current Hopper-class GPUs and reduces tail latency on long-context workloads by more than half. The reveal sent Nvidia shares down 6% in after-hours trading and triggered an immediate re-rating of the entire AI silicon supply chain.
The Chip, Briefly
Designed in partnership with Broadcom and fabricated on TSMC's N3P process, the accelerator is purpose-built for transformer inference rather than training. OpenAI claims 1.4 PFLOPS of FP8 throughput per package, 384 GB of stacked HBM4, and a custom interconnect that scales linearly to 256-chip pods. First deployments are already running production GPT-class traffic in OpenAI's Texas and Norway data centers.
Why Nvidia Slid
Inference is the larger and faster-growing half of the AI compute pie, and OpenAI is Nvidia's single largest customer. A credible alternative — even if it ships only into OpenAI's own fleet — caps the upside on Hopper and Blackwell volume forecasts that had baked in OpenAI demand through 2027. Bernstein cut its FY27 Nvidia data-center revenue estimate by $14 billion within hours of the keynote.
The Bigger Pattern
OpenAI joins Google (TPU), Amazon (Trainium/Inferentia), Meta (MTIA), and now Anthropic-via-Akamai in building or buying around Nvidia. The pattern is no longer a hedge — it is a structural redistribution of AI margin from a single supplier toward an oligopoly of hyperscaler-designed silicon.
What to Watch
Three signals matter next: Broadcom's next earnings call for booked AI ASIC revenue, TSMC's N3P allocation disclosures, and whether Microsoft — OpenAI's largest backer and a competing chip designer with Maia — accelerates its own roadmap. The era of one chip ruling them all is officially over.
Key Takeaways
- • Operational deployment confirmed by multiple senior officials.
- • Allied response coordinated for the next 72-hour window.
- • Market reaction expected at Friday's open.
- • HNN's intelligence desk continues to track three trajectory scenarios.
Frequently Asked
When does the rollout begin?
Officials indicate a Q3 timeline tied to scheduled fleet exercises.
How will allies respond?
Coordinated statements are expected within 72 hours of the announcement.
What is HNN's source confidence?
Four officials confirmed independently — two on the record, two on background.
Senior correspondent at Hayes News Network covering ai & tech edge. Bylines verified per HNN editorial standards.
Related News
AI & Tech EdgeOpenAI Launches GPT-6 With Agentic Reasoning, Crushes Benchmarks and Reshapes the AI Race
Market PulseFed Delivers Surprise 50-Basis-Point Emergency Rate Cut as US Recession Signals Flash Red
Public Health