OpenAI announced first results for Jalapeno, its in-house AI inference ASIC co-developed with Broadcom, claiming 1.5x-1.9x higher throughput per kilowatt and 1.7x-3.6x lower end-to-end latency than NVIDIA's GB200/GB300 rack systems. The chip carries 216GB of HBM4 and OpenAI plans to begin deploying it across its compute by the end of 2026.
OpenAI unveils Jalapeno, its first custom inference chip, co-built with Broadcom
Cheaper, faster inference silicon from a major lab signals downward pressure on API costs—AI product founders should factor potential price/latency improvements into 2027 unit-economics planning.
Source: CNBC
More that helps you.
Perplexity ships on-device Portable Computer agent in its Windows app
On September 15, 2026 Perplexity launched a local AI agent in its Windows app for Nvidia GeForce RTX and RTX PRO GPUs with at least 24GB of VRAM. The model, agent harness, orchestr…
Australia's Metacognition AI raises $10M for an AI operating-system layer
Australian lab Metacognition AI announced a roughly $10 million round on September 10, 2026, led by CSIRO-backed VC Main Sequence. It is building an operating-system layer that sit…
NVIDIA ships TensorRT Model Connect: Hugging Face checkpoint to C++ in two commands
On Aug 28, 2026, NVIDIA released TensorRT Model Connect, an open collection of reference implementations that turns a Hugging Face or local checkpoint into native C++ inference wit…
Factory raises $200M at $5B valuation for its AI coding agents
AI software-engineering startup Factory raised $200 million at a $5 billion valuation on Sept 16, 2026, roughly tripling the $1.5B valuation from its $150M Series C five months ear…
Get briefs like this tuned to you.
In the app, Founder Briefs are personalized to your country, industry and stage, and you can save the ones that matter.
See plans →