A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
What is the pricing for the FW-Kimi-K3 (by Fireworks) model in Azure AI Foundry?
Here's a revised description that conveys frustration while staying professional enough for a public Microsoft Q&A:
I'm using the FW-Kimi-K3 model offered through Azure's partnership with Fireworks AI and it is deployed in Azure AI Foundry (East US) under a Microsoft for Startups credit subscription. I cannot find published pricing for it anywhere: not on the Azure pricing page, the Foundry model catalog, or the Retail Prices API. This is frankly unacceptable for a service that is actively billing customers.
From my Cost Management records, usage appears to be billed at approximately $10 per million tokens, and this same flat rate is applied to every token type which is input, cached input, and output with no cache-read discount whatsoever. On July 28 alone, my deployment logged 18.55M input tokens and 272.1K output tokens, and approximately $188 in startup credits were debited. Charging the same rate for cached input tokens as for output tokens is unreasonable and inconsistent with how every other provider prices cache reads.
What makes this worse is that there was no way for me to know this rate in advance. The pricing is not documented, the meter names in Cost Management are generic (Model 12, Model 14, Model 15), and there is no indication that a flat $10/M rate would apply across all token types. As a startup relying on credits to evaluate models, I had no opportunity to estimate or control costs before they were incurred.
Could someone from Microsoft please explain:
- The official per-million-token price for
FW-Kimi-K3(input, output, and cached input separately, if applicable), and whether it is set by Fireworks or Azure. - Whether a cached-input discount exists for this model, and if not, why not.
- Where this pricing is published — or confirmation that it is currently undocumented and being billed without public disclosure.
- Why a flat rate is applied to all token types rather than differentiated input/output/cache pricing.
Startups on credit subscriptions deserve transparent, documented pricing before being billed. Without it, we cannot responsibly use or evaluate these models. Thank you.
Want me to tone the frustration up or down, or is this the right register?