LIVEBreaking
BREAKING: Pentagon confirms Pacific Fleet railgun rollout begins Q3MARKETS: S&P 500 futures up 0.8% in pre-market tradingGEOPOLITICS: EU and ASEAN sign new strategic partnership in GenevaTECH: Meta closes $14B acquisition of major AI labHEALTH: CDC issues new guidance on early-onset neurological screeningARCHIVE: HNN Investigates — The forgotten architects of Fatehpur SikriBREAKING: Pentagon confirms Pacific Fleet railgun rollout begins Q3MARKETS: S&P 500 futures up 0.8% in pre-market tradingGEOPOLITICS: EU and ASEAN sign new strategic partnership in GenevaTECH: Meta closes $14B acquisition of major AI labHEALTH: CDC issues new guidance on early-onset neurological screeningARCHIVE: HNN Investigates — The forgotten architects of Fatehpur Sikri
Advertisement
AI & Tech Edge

OpenAI Unveils Custom Inference Chip, Sending Nvidia Shares Tumbling

OpenAI took the wraps off its first in-house inference accelerator on Saturday, claiming a 4x cost-per-token advantage over current Hopper GPUs and confirming Broadcom and TSMC as silicon partners. Nvidia slid 6% in after-hours trading as analysts began rewriting their 2027 AI capex models in real time.

VH
By Victoria Hayes
May 10, 2026 · 7 min read read
Share
Glowing AI inference server racks with overlaid stock chart in a dark data center
Glowing AI inference server racks with overlaid stock chart in a dark data center — HNN / File Photo
Advertisement

SAN FRANCISCO — OpenAI on Saturday took the wraps off its first in-house inference accelerator, a chip the company says delivers a 4x cost-per-token advantage over current Hopper-class GPUs and reduces tail latency on long-context workloads by more than half. The reveal sent Nvidia shares down 6% in after-hours trading and triggered an immediate re-rating of the entire AI silicon supply chain.

The Chip, Briefly

Designed in partnership with Broadcom and fabricated on TSMC's N3P process, the accelerator is purpose-built for transformer inference rather than training. OpenAI claims 1.4 PFLOPS of FP8 throughput per package, 384 GB of stacked HBM4, and a custom interconnect that scales linearly to 256-chip pods. First deployments are already running production GPT-class traffic in OpenAI's Texas and Norway data centers.

Why Nvidia Slid

Inference is the larger and faster-growing half of the AI compute pie, and OpenAI is Nvidia's single largest customer. A credible alternative — even if it ships only into OpenAI's own fleet — caps the upside on Hopper and Blackwell volume forecasts that had baked in OpenAI demand through 2027. Bernstein cut its FY27 Nvidia data-center revenue estimate by $14 billion within hours of the keynote.

The Bigger Pattern

OpenAI joins Google (TPU), Amazon (Trainium/Inferentia), Meta (MTIA), and now Anthropic-via-Akamai in building or buying around Nvidia. The pattern is no longer a hedge — it is a structural redistribution of AI margin from a single supplier toward an oligopoly of hyperscaler-designed silicon.

What to Watch

Three signals matter next: Broadcom's next earnings call for booked AI ASIC revenue, TSMC's N3P allocation disclosures, and whether Microsoft — OpenAI's largest backer and a competing chip designer with Maia — accelerates its own roadmap. The era of one chip ruling them all is officially over.

Advertisement

Key Takeaways

  • • Operational deployment confirmed by multiple senior officials.
  • • Allied response coordinated for the next 72-hour window.
  • • Market reaction expected at Friday's open.
  • • HNN's intelligence desk continues to track three trajectory scenarios.

Frequently Asked

When does the rollout begin?

Officials indicate a Q3 timeline tied to scheduled fleet exercises.

How will allies respond?

Coordinated statements are expected within 72 hours of the announcement.

What is HNN's source confidence?

Four officials confirmed independently — two on the record, two on background.

VH
About the author
Victoria Hayes

Senior correspondent at Hayes News Network covering ai & tech edge. Bylines verified per HNN editorial standards.

Related News

Advertisement