Unbilled Pricing Discrepancy & Refund Request: DeepSeek V4 Flash "Cached Global Tokens" Charged at ~357x Published Rate

Jagannathan Veeraraghavan 20 Reputation points
2026-08-08T07:25:23.4833333+00:00

Environment Details:

  • Service: Azure AI Foundry (Serverless)

Model: DeepSeek V4 Flash

Deployment Dates: June 6–7, 2026

Subscription Billing Impact: ₹161,648.62 (~$1,900 USD)

Issue Context

During a two-day deployment of DeepSeek V4 Flash via Azure AI Foundry, our subscription incurred an unexpected charge of ₹161,648.62, with 98.4% of the total cost (₹158,990.13 / ~$1,684.33) originating from a single line item: "V4 Flash Cached Global Tokens".

When checking with Azure Support, we were informed that the usage report reflects an effective applied rate of $0.01 per K ($10.00 per 1M tokens) for cached tokens. Support advised us to post here on Microsoft Q&A for model-specific billing and pricing clarification.

The Discrepancy

We verified the math behind our invoice, but the applied cached token rate contradicts Microsoft's published pricing:

Input Tokens: 145,840.449K @ ~₹17.93/M ($0.19/M) = ₹2,615.62 (Verified & Matches Published Documentation)

Output Tokens: 890.597K @ ~₹48.14/M ($0.51/M) = ₹42.87 (Verified & Matches Published Documentation)

Cached Global Tokens: 168,432.896K @ $0.01 per K ($10.00 / 1M tokens) = ₹158,990.13 (~$1,684.33) (Major Discrepancy)

Comparison with Published Microsoft Documentation:

According to official Microsoft Tech Community documentation (Introducing DeepSeek V4 Flash and V4 Pro in Microsoft Foundry), the official rates are:

Input: $0.19 / 1M tokens

Output: $0.51 / 1M tokens

Cached Input: $0.028 / 1M tokens

The rate applied to our deployment ($10.00 / 1M tokens) is ~357× higher than Microsoft's officially published $0.028 / 1M tokens rate for Cached Input on DeepSeek V4 Flash.

Key Questions & Request for Resolution

Undocumented Rate Basis: Where did the $0.01 per K ($10.00 / 1M) effective rate originate? Was this rate ever publicly documented or disclosed prior to June 2026?

Meter Clarification: Is "V4 Flash Cached Global Tokens" intended to be the same meter as the documented "Cached Input" meter ($0.028 / 1M)?

Refund / Credit Request: Because an undisclosed rate was applied—representing a 357x markup over documented prices—we are seeking a billing adjustment/credit to recalculate this usage under the official $0.028 / 1M rate or credit the undocumented line item.

We appreciate any assistance or escalation from the Azure AI Foundry team to help resolve this discrepancy.

(Note: Screenshots of the invoice breakdown and official Microsoft pricing tables are attached below as reference.)deepseek-az-breakdown

WhatsApp Image 2026-07-23 at 20.27.36

Screenshot 2026-08-08 at 12.45.36 PM

Foundry Models
Foundry Models

A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference

0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.