Skip to content
Ops Log
Lab NoteSecurity & Infrastructure05 September 202611 min read

Azure Cost Management can return 429 with 597 of 600 tenant queries unused: the exhausted bucket is the client type

A Cost Management query returned 429 on 2026-08-14 while the response header for the documented per tenant quota still reported 597 of 600 hourly queries left. The exhausted bucket belonged to the caller's client type, and the same request body with a ClientType header returned 200 seconds later.

Cover image for Azure Cost Management can return 429 with 597 of 600 tenant queries unused: the exhausted bucket is the client type
MJ

A call to the Azure Cost Management Query API can return HTTP 429 while the quota Microsoft documents for that API is almost untouched. Measured on 2026-08-14 against one Azure subscription: the throttled response carried x-ms-ratelimit-microsoft.costmanagement-qpu-remaining: QueriesPerHour:597,QueriesPerMin:57,QueriesPer10Sec:11, against documented quotas of 600 per hour, 60 per minute and 12 per ten seconds. On the same response, x-ms-ratelimit-remaining-microsoft.costmanagement-clienttype-requests read DefaultQuota:0. Nothing about the subscription or the tenant was throttled. The empty bucket was attached to the caller's client type.

Last verified: 2026-09-05.

What Microsoft documents

The rate limits Microsoft publishes for this API sit on one page, Manage costs with automation, read 2026-09-05. It opens by saying that the model it is about to describe is not the only one: "In addition to the existing rate limiting processes, the Query API also limits processing based on the cost of API calls." It then documents that one model, query processing units, and it is explicit about the scope they attach to:

The following quotas are configured per tenant. Requests are throttled when any of the following quotas are exhausted.

The three quotas it lists are 12 QPU per 10 seconds, 60 QPU per 1 min and 600 QPU per 1 hour, and it adds that the quotas "maybe be changed as needed and more quotas can be added", spelling included. The same page names three response headers, all of them QPU headers, and says of the retry one: "Indicates the time to back-off in seconds. When a request is throttled with 429, back off for the time specified in this header before retrying the request."

The Query API reference, api-version 2026-06-01, read the same day, names a different header again. Its ErrorResponse definition says a 429 means the request "is throttled" and directs the caller to the time specified in the x-ms-ratelimit-microsoft.consumption-retry-after header. The resource provider in that name is consumption, not costmanagement.

Neither page mentions a client type, a ClientType request header, or a quota bucket that several callers can share. That absence is the finding, and this note does not claim otherwise.

Where the ClientType header is described

The only Microsoft-hosted descriptions found on 2026-09-05 are posts on Microsoft Q and A, which is a support forum with moderator participation rather than reference documentation. They are quoted as what they are, and none is a published limit.

On a thread from August 2023, a moderator badged Microsoft External Staff told a caller who was making few requests that the throttling came from a ClientType filter "which allows 2000 calls per minute per ClientType", and added the sentence that explains the reading above: "since you are not providing a ClientType in the request headers, you will share the number with ALL customers that don't provide a ClientType". The same answer, dated 3 August 2023, listed four retry headers to look for on a 429, one of them ending costmanagement-client-retry-after. The original poster confirmed on 17 August 2023 that adding the header solved it.

Two months later, on 17 October 2023, another reader asked the same moderator on that thread to "link to the documentation showing the use of the 'ClientType' header? I can't seem to find it", followed the next day by asking whether the answer was that no such documentation exists, and was pointed at the feedback control on the Cost Management REST reference. That exchange is the strongest support for treating the header as undocumented rather than as something a search missed.

The same moderator, in a comment dated 25 October 2024 on a Grafana thread whose only posted answer is by a volunteer moderator, gave the mechanics: "Add a ClientType request header. The client request header is a http header that you can send along with the Authorization header. The key would be ClientType, the value could be any string". The first recommendation in that same comment is spacing rather than a header, "Add 20 seconds wait time between each API call", worked through on the page to at most three calls per minute, and it names the buckets that spacing protects: "In that way, you can avoid entity, tenant and QPU rate limit." Entity and tenant are two of the three counters observed below, and neither appears on either reference page. A third thread, answered 16 July 2025, tells a Go SDK caller to set the value through the client's ApplicationID telemetry option, and that answer carries a Microsoft External Staff badge without the moderator one.

The same pages carry reports that the header changed nothing, and an article that quoted only the confirmations would be quoting half a thread. On thread 1340993 a second caller wrote on 24 November 2023 that "I tried it many times in last 3 days, this header does not work for me", and reported on 30 November 2023 that what separated success from failure in that case was the token rather than the header. A further answer on the same thread, dated 23 February 2024, credits a different header entirely: "Adding this header solved the issue: X-Ms-Command-Name: CostAnalysis Adding the ClientType header does not make any difference". The Grafana thread opens from that same position, because its asker was already sending both a ClientType and an x-ms-command-name header and was throttled anyway.

Why the documented headers point the wrong way

An operator who reads the documented page and then reads the response is led, fairly, to the wrong conclusion. The QPU headers are the ones Microsoft names, they are present on the throttled response, and they report room: 597 of 600 for the hour, 57 of 60 for the minute, 11 of 12 for the ten seconds. Nothing in the QPU model can produce a 429 at those numbers. The page does say other rate limiting processes exist, and it names none, states no quota for them and documents no header that reports one, so the readings left are that Cost Management is intermittently unavailable, or that the retry logic is too timid.

Both readings lead to retrying, and retrying is the one move that makes this worse. The measured x-ms-ratelimit-microsoft.costmanagement-clienttype-retry-after grew from 4 seconds to 38 seconds across attempts on 2026-08-14, so a fixed interval retry loop lengthens its own penalty while producing no data.

The cost lands one layer up. A collector that catches the 429 and carries on renders a subscription that is billing as zero, or an estate with nothing in it, or a month that improved. That output is indistinguishable from good news. It is the same degraded read trap as a permissions error that returns fewer rows: fewer rows is less evidence, not less cost, and a throttled read has to set a health flag to false and produce a partial result that says so on its face.

Test conditions

Run on 2026-08-14 against one Azure subscription under a Microsoft Customer Agreement billing profile, as a signed in administrator through az rest to Azure Resource Manager, Cost Management query api-version 2025-03-01. Three single shot calls seconds apart, same bearer token, same subscription, same request body, no retry logic and no delay policy, so the header was the only variable. The endpoint is a POST that reads and creates nothing. The vendor pages were re-read on 2026-09-05 and every sentence quoted here was confirmed present.

http
# Read-only. Send the same body twice: once without the ClientType header,
# once with. api-version 2026-06-01 is what the reference documents on 2026-09-05.

POST https://management.azure.com/subscriptions/{subscriptionId}/providers/Microsoft.CostManagement/query?api-version=2026-06-01
Authorization: Bearer {token for an identity holding the built-in Reader role}
Content-Type: application/json
ClientType: Contoso-CostCollector

{
  "type": "ActualCost",
  "timeframe": "MonthToDate",
  "dataset": {
    "granularity": "None",
    "aggregation": { "totalCost": { "name": "Cost", "function": "Sum" } }
  }
}

# Read these headers on EVERY response, 200 and 429 alike:
#   x-ms-ratelimit-microsoft.costmanagement-qpu-remaining
#   x-ms-ratelimit-remaining-microsoft.costmanagement-clienttype-requests
#   x-ms-ratelimit-remaining-microsoft.costmanagement-tenant-requests
#   x-ms-ratelimit-microsoft.costmanagement-clienttype-retry-after
#
# A 429 whose qpu-remaining is high while clienttype-requests reads 0 is the
# case in this note.

Results

Call A, no ClientType header: 429. Call B, a ClientType header naming the collector, everything else identical: 200. Call C, header removed again: 429. The header is the cause rather than elapsed time, because time passing between A and C would have relieved a time based quota instead of restoring the failure.

The throttled responses carried x-ms-ratelimit-remaining-microsoft.costmanagement-clienttype-requests: DefaultQuota:0 beside x-ms-ratelimit-remaining-microsoft.costmanagement-entity-requests: DefaultQuota:2 and x-ms-ratelimit-remaining-microsoft.costmanagement-tenant-requests: DefaultQuota:17. Three buckets, three separate readings, one of them empty. None of those three is the documented QPU quota, whose own header on the same response reported a tenant with room while the client identity had none.

One check does not depend on trusting this measurement. The question that opens thread 1340993, posted on 3 August 2023 by an unrelated caller, pastes three throttled responses of its own, and each one shows costmanagement-clienttype-requests: DefaultQuota:0 next to entity readings of 3, 2 and 0 and tenant readings of 19, 18 and 14. Same header names, same shape, a different tenant three years earlier: the client bucket empty while the tenant bucket still has room.

Failures

Four checks against the documentation failed. First, the documented QPU headers did not explain the observed 429, and the page's one acknowledgement that other rate limiting processes exist is all it says about them. Second, the header that did arrive, x-ms-ratelimit-microsoft.costmanagement-clienttype-retry-after, appears on neither reference page, though it is printed in the 2023 forum question quoted above. Third, the two reference pages name two different retry headers for the same status code on the same API. Fourth, the automation page states its quotas in QPU while the header it documents reports them as queries, so a reader has to assume the two words mean one thing.

Source, read 2026-09-05What it says is throttledThe retry header it names
Manage costs with automationQPU quotas, configured per tenantcostmanagement-qpu-retry-after
Query API reference, 2026-06-01429 TooManyRequests, no dimension namedconsumption-retry-after
Microsoft Q and A answer, 3 August 20232000 calls per minute per ClientTypecostmanagement-client-retry-after, plus three others
Observed on the wire, 2026-08-14clienttype bucket at DefaultQuota:0costmanagement-clienttype-retry-after

Four names for one mechanism, and the two reference pages carry neither of the names that arrive on a real 429. Plan against the wire: read every x-ms-ratelimit- header present on the response rather than the ones a page said to expect, and branch on whichever remaining count reads zero.

Limitations

One subscription, one tenant, one day, one client. The comparison establishes that the header changed the outcome on 2026-08-14. It does not establish that a distinct ClientType prevents a 429, and it cannot, because a dedicated bucket is still a bucket with a limit. A caller that exceeds its own allowance gets the same status code with a different reading in the same header.

The cited threads carry the same limit from the other direction. Two callers there report no change after adding the header, one crediting a different header and the other a different token, and the Grafana thread is opened by a caller who was already sending it. A distinct ClientType stops one estate sharing a bucket with strangers. It does not make a caller that queries too often stop being throttled.

The 2000 calls per minute figure is a moderator's answer on a support thread from August 2023, not a published limit, and Microsoft's own quota page reserves the right to change quotas and add more. Treat the figure as an explanation of the mechanism rather than as a number to size a pipeline against.

The measurement says nothing about how the shared bucket is keyed beyond the wording quoted above, nothing about which callers were in it, and nothing about other Cost Management endpoints. The quotas and the ClientType behaviour described here belong to the Query API; the Cost Details report path is asynchronous and carries its own guidance, which tells the caller to use the retry-after header in the response to decide when to poll again.

Reading a header is not a substitute for calling less. Microsoft states that Cost Management data "is refreshed every four hours as new usage data is received from Azure resource providers" and that calling more frequently "creates increased load", so a collector polling faster pays for identical rows twice.

What to change today

Send a distinct, stable ClientType on every Cost Management call, named for the calling system rather than for a run. Honour the clienttype retry header instead of a fixed interval. Log the remaining counts per call. And make a 429 set the health flag to false, so the report it feeds is marked partial rather than published as a number.

The Azure half of the SEAWALL evidence pack is generated from ITSailor's own Azure subscription and carries the health flag this note describes, including what the pack refuses to state when a read comes back degraded.

Sources and further reading

Was this field note useful?
Use the evidence

Test the same boundary in your environment.

Use a focused diagnostic to compare the lab result with the controls and constraints in your own environment.

Choose a diagnostic