Forum Discussion
Intune On-Demand Proactive Remediation API Reliability for Large-Scale Usage
Your 18–19 successful executions out of 20 requests warrant investigation, but an HTTP 204 response is not endpoint execution evidence. Microsoft documents that repeated Run remediation actions for the same device can overwrite each other, and delivery requires working Intune and Windows Push Notification Service connectivity at submission time. A fixed 30-second delay therefore does not guarantee delivery. Serialize actions per device, verify its Remediations monitoring status, and correlate timestamps with Intune Management Extension logs before retrying. For HTTP 429 responses, honor Retry-After; use exponential backoff when that header is absent. Across devices, pilot bounded concurrency and measure confirmed execution rather than request acceptance. Microsoft supports Intune beta APIs, but this on-demand feature remains documented as preview. The cited documentation provides neither a fleet-scale delivery guarantee nor a GA date. Escalate the reproducible misses with request IDs and logs before treating it as a guaranteed production orchestrator
You can post/reply with something like this:
Hi Jamony,
Thank you for the response.
We are trying to understand the complete behavior of the On-Demand Remediation API before recommending or deploying this approach for customers at scale.
In our testing, we observed that remediation requests triggered manually from the Intune portal appear to be more reliable. However, our understanding is that the portal is also using the same backend API. Because of that, we would like to better understand the underlying architecture and delivery mechanism.
Some of the areas we are trying to clarify are:
- What is the recommended retry strategy for failed or missed executions?
- If a request returns HTTP 204, what additional validation should be performed to confirm that the remediation was actually queued and executed on the endpoint?
- How exactly are remediation requests processed in the backend? Is there a queueing mechanism, and if so, are there any documented limits or retention periods?
- When multiple requests are submitted for the same device, how does the overwrite behavior work? Does the latest request replace all previous pending requests or only specific ones?
- Is there any difference between requests initiated from the Intune portal versus those triggered through Microsoft Graph APIs?
- Are there known throttling thresholds, concurrency limits, or scaling recommendations for larger environments?
- Is delivery dependent on the device maintaining active Intune Management Extension (IME) and Windows Push Notification Service (WNS) connectivity at the exact submission time?
- Are there any telemetry fields, request identifiers, or logs that can be used to reliably correlate API requests with actual endpoint execution?
- Has Microsoft published any guidance on expected success rates, supported scale, or production best practices for this feature?
- Are there any plans for General Availability (GA), or is the preview behavior expected to change significantly?
Also, if you are aware of any detailed architecture documentation, engineering blogs, whitepapers, internal presentations shared publicly, or deeper technical references beyond the currently available documentation, could you please share them?
Our goal is to fully understand the delivery flow, queueing behavior, retry handling, overwrite scenarios, and monitoring capabilities so that we can design a reliable and scalable solution for customer environments.
Thank you for your help. We appreciate any additional insights you can provide.