A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Unbilled Pricing Discrepancy & Refund Request: DeepSeek V4 Flash "Cached Global Tokens" Charged at ~357x Published Rate
Environment Details:
- Service: Azure AI Foundry (Serverless)
Model: DeepSeek V4 Flash
Deployment Dates: June 6–7, 2026
Subscription Billing Impact: ₹161,648.62 (~$1,900 USD)
Issue Context
During a two-day deployment of DeepSeek V4 Flash via Azure AI Foundry, our subscription incurred an unexpected charge of ₹161,648.62, with 98.4% of the total cost (₹158,990.13 / ~$1,684.33) originating from a single line item: "V4 Flash Cached Global Tokens".
When checking with Azure Support, we were informed that the usage report reflects an effective applied rate of $0.01 per K ($10.00 per 1M tokens) for cached tokens. Support advised us to post here on Microsoft Q&A for model-specific billing and pricing clarification.
The Discrepancy
We verified the math behind our invoice, but the applied cached token rate contradicts Microsoft's published pricing:
Input Tokens: 145,840.449K @ ~₹17.93/M ($0.19/M) = ₹2,615.62 (Verified & Matches Published Documentation)
Output Tokens: 890.597K @ ~₹48.14/M ($0.51/M) = ₹42.87 (Verified & Matches Published Documentation)
Cached Global Tokens: 168,432.896K @ $0.01 per K ($10.00 / 1M tokens) = ₹158,990.13 (~$1,684.33) (Major Discrepancy)
Comparison with Published Microsoft Documentation:
According to official Microsoft Tech Community documentation (Introducing DeepSeek V4 Flash and V4 Pro in Microsoft Foundry), the official rates are:
Input: $0.19 / 1M tokens
Output: $0.51 / 1M tokens
Cached Input: $0.028 / 1M tokens
The rate applied to our deployment ($10.00 / 1M tokens) is ~357× higher than Microsoft's officially published $0.028 / 1M tokens rate for Cached Input on DeepSeek V4 Flash.
Key Questions & Request for Resolution
Undocumented Rate Basis: Where did the $0.01 per K ($10.00 / 1M) effective rate originate? Was this rate ever publicly documented or disclosed prior to June 2026?
Meter Clarification: Is "V4 Flash Cached Global Tokens" intended to be the same meter as the documented "Cached Input" meter ($0.028 / 1M)?
Refund / Credit Request: Because an undisclosed rate was applied—representing a 357x markup over documented prices—we are seeking a billing adjustment/credit to recalculate this usage under the official $0.028 / 1M rate or credit the undocumented line item.
We appreciate any assistance or escalation from the Azure AI Foundry team to help resolve this discrepancy.
(Note: Screenshots of the invoice breakdown and official Microsoft pricing tables are attached below as reference.)