DeepSeek released V4.1 Flash on September 10, 2026, priced at $0.15 per million uncached input tokens and $0.60 per million output tokens off-peak (double during 01:00-04:00 and 06:00-10:00 UTC weekdays). Weights ship under an MIT license for self-hosting, and V4.1 Flash takes over V4 Pro traffic on September 14.
If you run LLM features, benchmark V4.1 Flash against your current model; the off-peak rates and open weights can cut inference costs sharply for high-volume or self-hosted workloads.
Source: The Rundown AI
More that helps you.
OpenAI ships GPT-6 Astra with 1M+ context at $10/$50 per million tokens
OpenAI released GPT-6 Astra (API id gpt-6-astra) on Sept 3, 2026, priced at $10 per million input tokens and $50 per million output, with cached input at $1. It carries an ~1.1M-to…
Google's Gemini 3.8 Flash launches at $0.75/$3.75 per million, doubling in 2027
Google released Gemini 3.8 Flash on Sept 2, 2026 with a 1,048,576-token context window, priced at $0.75 input and $3.75 output per million tokens through Dec 31, 2026. The rate is…
Meta ships Muse Spark 1.3, cutting tool calls ~20% and tokens ~25%
Meta released Muse Spark 1.3 on Sept 2, 2026 at $1.25 input and $4.25 output per million tokens with a 1,048,576-token context window, claiming roughly 20% fewer tool calls and 25%…
IBM releases open Granite 4.2 models (3B–30B) under Apache 2.0
On Aug 26, 2026 IBM released Granite 4.2, a family of open models in 3B, 8B, and 30B parameter sizes with a 128,000-token context window under the permissive Apache 2.0 license. Th…
Get briefs like this tuned to you.
In the app, Founder Briefs are personalized to your country, industry and stage, and you can save the ones that matter.
See plans →