An AI inference price hike is coming from the one company nobody expected to raise prices. DeepSeek is the Chinese lab that triggered the global collapse in token costs. Last Thursday it told developers its API prices are going up, and that the increase could be substantial. No new rate card. No effective date. Just a warning. For two years, every AI budget in software has been built on the assumption that tokens only get cheaper. That assumption cracked this month, and HR software buyers are downstream of it.
The AI Inference Price Hike Came From the Company That Started the Price War
DeepSeek published a notice on August 6 saying it plans to raise API prices soon, warning users the change could be significant. The company has not published a new price schedule, so the increase is planned rather than live. (Source: TechNode)
The timing is what makes it interesting. A week earlier, DeepSeek had shipped a coding model at a startling discount. Axios put it at roughly 99 percent against a leading US frontier model. In recent weeks the company also started charging surcharges for API calls during peak hours. So the same company was cutting headline prices and quietly adding usage penalties at the same time. (Source: Semafor)
Demand is the obvious culprit. DeepSeek’s V4 Flash topped OpenRouter’s weekly model ranking for July 27 to August 2 with 7.22 trillion tokens processed. (Source: TechNode) Coding tool OpenCode separately reported roughly 8 trillion tokens in a single day on August 1.
Meanwhile, the company reopened a funding round of about $8 billion at a valuation near $74 billion. Part of that money is earmarked for data centres. (Source: Bloomberg) Raising capital for compute while raising prices tells you the unit economics were never comfortable.
Why This AI Inference Price Hike Reaches Your HR Stack
Here is the part most HR buyers miss. You do not pay for tokens, but your vendors do, and they priced their AI features when tokens were close to free.
Rippling gave the clearest public accounting of this. In March, CFO Adam Swiecicki brought a number to the executive team. The company was on pace to spend 40 percent of its R&D headcount budget on AI tokens. Spending was growing about 80 percent month over month. Roughly 10 to 15 percent of employees drove about 60 percent of that bill. One engineer alone was running $50,000 a month. (Source: TechCrunch)
That is an HR software company, with real engineering discipline, nearly matching its R&D payroll with an inference bill. Now add an AI inference price hike on top of it.
Rippling’s fix is instructive. It built an internal console to track spend per employee against output and route routine work to cheaper models. Token use barely moved. April ran 605 billion tokens, July ran 600 billion. Yet the July bill came to 37 percent of April’s cost, and the spend fell to about 15 percent of the R&D headcount budget. (Source: Rippling)
In other words, the savings came from routing, not from restraint. That lever exists because cheap models exist. If the cheap tier reprices, the lever gets shorter for everyone building AI agents for HR workflows.
Under the Hood: What Actually Broke the Cheap-Token Model
Inference is not software margin. Every generated token consumes GPU time, memory bandwidth, and power. Selling below cost works as a customer acquisition strategy, but only while volume stays modest.
Agentic workloads broke that assumption. A chatbot answer costs a few thousand tokens. A coding agent or a document-processing agent burns millions, because it loops, retries, reads long files, and calls tools. Consequently, the cheapest models attract exactly the workloads that consume the most compute.
HR workloads sit squarely in that expensive category. Screening a thousand applications, reading long policy documents, reconciling payroll exceptions across countries, summarising every performance review cycle. Each of those is a long-context, high-repetition job. Therefore the AI features your HR team likes most are usually the ones costing your vendor the most to run.
Peak-hour surcharges are the tell. When a provider prices by time of day, it has run out of spare capacity, not out of customers. An AI inference price hike is what follows once the queue stops clearing.
The pressure is not one-directional, though. Competitors are still cutting headline prices. OpenAI reduced the cost of a new model this month, and Meta launched a coding model with cost as the pitch. So the market is splitting. Sticker prices keep falling on new releases, while real, sustained, high-volume usage gets more expensive through surcharges, rate limits, and repricing. Averages will look calm. Your invoice may not.
What HR Leaders Should Do About Rising Inference Costs on Monday
First, find out whether your HR vendors expose AI usage limits in your contract. Many bundle AI features into a per-employee price with a fair-use clause that lets them reprice or throttle later. Ask directly what happens if their model costs double, because that clause is where the risk lives.
Second, ask which model tier powers each AI feature you actually use. Resume parsing, leave classification, and policy question answering do not need a frontier model. If your vendor cannot answer, it probably has not measured, and unmeasured spend gets passed on.
Third, measure your own internal AI spend the way Rippling did. Track cost against output per team, not per licence. A concentrated few users usually drive most of the bill. That is a manageable problem once you can see it, and an invisible one until then.
Fourth, put renewal dates and AI pricing terms in the same calendar. Vendor pricing pages will change quietly over the next two quarters. Teams reviewing their AI tools for HR now will negotiate from a better position than teams that discover the change at renewal.
Finally, be careful about which workflows you make dependent on the cheapest tier. AI payroll automation that only works at $0.14 per million tokens is a business process with a price-risk attached.
Asanify builds HR, payroll, and EOR software for teams running across borders. We publish our pricing openly, because buyers deserve to know what a platform costs before the AI bill arrives. If you are auditing your HR stack this quarter, cost predictability belongs on the checklist next to features.
AI Inference Price Hike: FAQ
What is the AI inference price hike and when does it take effect?
DeepSeek said on August 6, 2026 that it plans to raise its API prices soon and warned the increase could be significant. The company has not published new rates or an effective date, so the change is announced rather than live. Developers should treat it as a signal about direction rather than a specific number.
Will other AI providers raise prices too?
Not uniformly. Headline prices on newly released models are still falling, with OpenAI and Meta both competing on cost this month. However, sustained high-volume usage is getting more expensive through peak-hour surcharges, rate limits, and repricing of older tiers.
How does this affect HR software buyers?
HR vendors priced their AI features when tokens were near free, and many contracts include fair-use clauses that allow later repricing or throttling. Ask your vendor which model tier powers each AI feature and what happens to your price if their inference costs rise. Track your own AI spend by team and output, because concentrated usage usually drives most of the bill.
Not to be considered as tax, legal, financial or HR advice. Regulations change over time so please consult a lawyer, accountant or Labour Law expert for specific guidance.
