The price on the pricing page is bait; the price you pay is set later, by the vendor, on a clock you do not hold.
DeepSeek moved its flagship model to general availability on August 12, the signal a lab sends when it wants you to build on something in production. Four days was all the reassurance lasted. In the same announcement, the lab said peak-hour output pricing on DeepSeek-V4-Pro would rise to $3.96 per million tokens from a flat $0.87, effective August 16. General availability usually means stability. Here it meant the introductory rate was over.
That is the week in one detail. Prices moved a lot, and the headline numbers fell, but every discount carried an expiry the buyer did not control.
Anyone picking the model behind a product this week had a genuinely cheaper menu to shop. The question is whether the price you evaluate is the price you will actually pay. For buyers, that gap is a dependency to test before the next planning cycle. For anyone who sells software, evals, or advisory work, it is a concrete thing to package. Both sides run through the same fact: a token price is a variable a supplier meters, not a rate you locked.
The discounts arrive pre-scheduled to end. Google put out Gemini 3.7 Flash on August 13, roughly three weeks after Gemini 3.6 Flash landed flat, and cut the price in half to pull developers back, to $0.75 per million input tokens and $3.75 output. Read the fine print and the discount expires December 31 and doubles on January 1. The cut and its reversal shipped together. A team wiring in Flash-tier calls today at $0.75 is building a budget on a number with a known expiry, four months out.
Grok 4.6, released August 12 by Elon Musk's AI company, tells the other half of the story. It scored 61 on the Artificial Analysis Index, an aggregate benchmark of model quality, level with OpenAI's best current model, and charges $2 per million tokens in and $6 out. Frontier-level quality at a discount is what turns the default-model choice into a monthly negotiation rather than a settled bet. That pressure is real and it helps buyers. It is also exactly the dynamic that makes a low launch price a customer-acquisition instrument: get wired in first, reprice once the switching cost is paid.
Why the floor keeps dropping, and why it will not stay down. The reason competent models keep getting cheaper is that competent models keep getting given away. Alibaba published the open weights for Qwen3.8-Max on August 13, a 2.4-trillion-parameter model it had sold only through an API ten days earlier. Nvidia is training Nemotron 4, an open family aimed at a trillion parameters, free to run. Meta released Muse Glimmer, a 30-billion-parameter agent that fits on a single 24GB GPU under an Apache license. When the mid-tier is commoditized by free weights, paid vendors cannot win on the sticker price. They win the default slot with a loss leader, then move the meter once the integration is done. DeepSeek's four-day repricing is that logic with the timeline compressed.
Hold two of this week's claims at arm's length. DeepSeek says its top tier clears 80 percent on SWE-bench Verified, a real-bug-fixing test, and no outside lab has checked it. Google's "half price" is accurate only until December 31. A benchmark you cannot reproduce and a price with an expiry are both attributed claims, not settled facts, and a unit-economics model built on either is building on sand.
Who this puts on the clock. The leverage sits with whoever captured the default slot and holds the repricing schedule. The exposure sits with any team that hardcoded one model at its launch rate and wrote the business case around it. That exposure is not abstract. A 4.5x jump in peak output cost, the size of DeepSeek's move, can flip a high-volume agent feature from profitable to underwater between one billing cycle and the next, and the buildout underneath these labs guarantees the pressure runs one way. Anthropic alone committed more than $60 billion to compute in three months, including a $9.1 billion, 20-year lease signed August 11. Nvidia lined up more than $500 billion of outside financing for data centers built around its chips. Those bills come due on someone's meter, and the meter is the API.
So the model choice is not a procurement decision you close at a fixed price. It is a metered dependency you re-open on the vendor's schedule.
For buyers and operators, stop pricing your AI features on the number you signed up at. Rerun the unit economics on the scheduled rate: Gemini Flash at its January 1 price, DeepSeek at the August 16 rate, not the one that pulled you in. Then wire one fallback model into your highest-volume path and keep it tested, so a repricing event is a config change instead of a rebuild. Run a hundred real calls through your current model and one alternate, compare cost, error rate, and latency, and you will know your switching cost before a vendor forces you to learn it.
For sellers, consultants, and software teams, there is a thirty-day model-cost exposure audit to package here. Map every model call in a client's stack, price each against its scheduled rate rather than today's, flag the calls sitting on an intro discount, and add fallback routing so a price move is a switch and not a fire drill. The proof artifact is a one-page exposure map with the failure cases attached. This week's repricing gives you the reason a buyer signs it: the risk stopped being hypothetical the day general availability and a 4.5x increase arrived in the same sentence.
The old assumption worth dropping is that AI compute is on a permanent glide toward zero, and that you can wire in the cheapest option and watch the bill shrink. The floor for open, self-hosted capability is falling. The price of the hosted convenience you actually build on is set deal by deal, by the party holding the clock. Treat the number on the pricing page as an opening offer, keep a second model warm, and the next repricing email becomes a decision you make rather than one made for you.
The week in one line: A hosted model's token price is a variable your vendor reprices on its own schedule, so evaluate it at the scheduled rate and keep a tested fallback, not a locked-in bet.
Sources this week: DeepSeek V4-Pro repricing, Gemini 3.7 Flash half-price, Grok 4.6 undercuts the frontier, Alibaba open-sources Qwen3.8-Max, Nvidia Nemotron 4, Meta Muse Glimmer, Anthropic's $9.1B compute lease, Nvidia's $500B financing