Webhooks vs. scheduled polling for real-time account data sync reliability

We’re building a real-time account synchronization between Dynamics 365 Sales 9.2 and our enterprise data warehouse. Currently using scheduled polling every 5 minutes via the Web API with change tracking, but considering moving to webhooks for true real-time updates. Our data team wants account changes reflected within 30 seconds for downstream analytics and reporting dashboards.

The polling approach is straightforward and we have good visibility into sync status, but it feels inefficient making API calls every 5 minutes when changes might be infrequent. Webhooks seem more elegant - push notifications when accounts actually change - but I’m concerned about reliability. What happens if our webhook endpoint is temporarily unavailable? How do you handle webhook delivery failures and ensure no data loss?

Looking for real-world experiences with both approaches, especially around failure scenarios and operational complexity.

Webhooks in Dynamics 365 Sales use the Service Endpoint + Plugin Registration Tool (or Power Platform CLI) pipeline, not a pure REST webhook subscription. Understanding that architecture is essential before you swap out polling.


How D365 Webhook Delivery Actually Works

D365 posts to your endpoint synchronously during the platform event pipeline. If your endpoint returns anything other than 2xx within the timeout window (~60 seconds, verify in your version), the platform marks the step as failed. Critically: D365 does not have a built-in retry queue with exponential backoff. Failed webhook calls surface as Plugin Trace Log entries and System Job failures, but the event is not automatically replayed. This is your primary reliability risk.

Webhook Registration (Plugin Registration Tool)

Service Endpoint:
  - Contract: WebHook
  - URL: https://your-middleware.example.com/d365-hook/account
  - Auth: HttpHeader  →  X-API-Key: <secret>

Step Registration:
  - Message: Update
  - Primary Entity: account
  - Filtering Attributes: name, telephone1, address1_city, <your critical fields>
  - Execution Mode: Asynchronous  ← use this, not Synchronous
  - Deployment: Server

Running the step Asynchronous offloads execution to the async service, decouples it from the user transaction, and allows the System Job to surface failures—but still does not auto-retry.

Handling Delivery Failures

Your middleware/endpoint layer must carry the reliability burden:

  • Implement an idempotent receiver keyed on x-ms-dynamics-request-id (header injected by D365) to safely handle duplicates if you build your own retry.
  • Front your endpoint with a durable queue (Azure Service Bus, SQS, etc.). D365 posts → queue ingests → downstream consumer processes. Queue handles your endpoint unavailability window cleanly.
  • Enable dead-letter queues and alert on them. This is your data-loss safety net.
  • Log organizationid + entityid + changedfields immediately on receipt before any processing.

Hybrid Pattern for Your 30-Second SLA

Neither pure webhooks nor pure polling alone is optimal here. Use both:

Layer Mechanism Purpose
Primary Webhooks (async step) Sub-30s delivery on happy path
Recovery Change tracking poll (15–30 min interval) Catch missed events, gap-fill
Validation modifiedon watermark comparison Detect drift between DW and D365

The OData change tracking endpoint ($select=name,modifiedon&$deltatoken=...) remains your consistency guarantee. Reduce polling to 15–30 minutes—you’re not relying on it for latency, only for gap recovery.

Version Compatibility Note

Webhook behavior, async service throughput limits, and System Job retention policies vary between on-premises 9.x and online. Verify your async service concurrency limits and System Job purge schedules in your specific deployment—online environments purge completed/failed jobs on a rolling schedule that can affect your failure audit window.

Operational Reality

Webhook-only for a DW sync with no retry infrastructure is fragile. The durable queue pattern eliminates the false choice between “elegant but risky” and “polling forever.”


This draft is based on general Microsoft Dynamics 365 Sales knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.

We migrated from polling to webhooks last year and it’s been a mixed experience. The latency improvement is real - we get updates within seconds instead of minutes. However, webhook reliability requires significant infrastructure investment. You need a message queue to buffer incoming webhooks, retry logic for processing failures, and monitoring to detect when webhooks stop arriving. We use Azure Service Bus as our webhook endpoint, which gives us durability and replay capabilities if our processing pipeline has issues.

One major gotcha with D365 webhooks: they don’t have built-in retry logic. If your endpoint returns an error or times out, the webhook is lost. You need to implement your own reconciliation process - typically a periodic polling job that checks for missed updates by comparing change tracking timestamps. So you end up maintaining both webhook and polling code anyway. For account data specifically, consider whether you really need sub-minute latency. Often business requirements say “real-time” but the actual use case works fine with 5-10 minute delays.

The point about needing polling as a safety net is concerning - that defeats the main benefit of webhooks. How do you balance the webhook processing with the reconciliation polling? If you poll every hour to catch missed webhooks, you’ve introduced up to 60 minutes of potential data lag for failure scenarios. And monitoring becomes more complex since you need to detect both webhook delivery issues and polling failures.

The hybrid approach is actually standard practice for mission-critical integrations. We run webhooks for the happy path and polling every 15 minutes as a safety net. The polling job only processes changes that aren’t already in our system based on ModifiedOn timestamps, so overhead is minimal. We also set up alerts if the gap between last webhook and last poll exceeds our SLA threshold. This architecture has proven very reliable - we get the latency benefits of webhooks while maintaining data consistency guarantees.

Consider API throttling in your decision too. With polling, you control the rate and can stay well under service protection limits. With webhooks, you’re at the mercy of user activity patterns. If someone bulk updates 1000 accounts, you’ll get 1000 webhook calls in rapid succession. Your endpoint needs to handle these bursts gracefully, probably by queuing them for async processing. This adds complexity but is manageable with proper architecture - Azure Functions with Service Bus triggers works well for this pattern.

Let me share our complete architecture that addresses all these concerns:

Webhook Configuration: D365 webhooks (called Service Endpoints) are configured at the entity level with specific message types (Create, Update, Delete). The critical setup decisions:

  1. Endpoint resilience: Don’t point webhooks directly at your processing logic. Use a durable message queue (Azure Service Bus, AWS SQS) as the webhook target. This decouples delivery from processing.

  2. Authentication: Webhooks include a shared secret for validation. Verify this on every incoming request to prevent spoofing.

  3. Payload design: Configure whether you want full entity payload or just change notification. For account data, we use change notification only - it includes the account ID and change type, then we fetch full details via API. This reduces webhook payload size and gives us current data even if processing is delayed.

Polling Intervals: Our hybrid strategy uses two polling mechanisms:

Primary reconciliation: Every 15 minutes, query accounts modified since last successful sync using change tracking. This catches any webhooks that failed delivery or got lost.

Deep reconciliation: Daily full sync comparing all account records between systems. This catches edge cases like webhooks that arrived but processing failed silently.

The 15-minute interval gives us a maximum 15-minute lag in failure scenarios, which meets our business SLA. Adjust based on your latency requirements.

API Throttling: This is crucial for webhook-based architectures. Our approach:

  1. Queue all incoming webhooks in Service Bus without processing
  2. Consumer processes queue with controlled concurrency (10 parallel threads)
  3. Implement exponential backoff when hitting throttling limits
  4. Use batch API calls where possible to reduce total request count

For bulk update scenarios (1000+ accounts), the queue absorbs the burst and we process at a sustainable rate. Monitor queue depth to detect processing bottlenecks.

Error Recovery: Multi-layered approach:

Webhook delivery failures: Detected by comparing webhook timestamp in Service Bus to change tracking queries. If we find modified accounts not in our recent webhook history, they’re queued for processing.

Processing failures: Failed messages go to dead letter queue after 3 retry attempts. Separate job processes dead letter queue with manual review for persistent failures.

Data validation failures: Logged separately with alerts. Often indicate schema changes or data quality issues that need investigation.

Operational Metrics: Key monitoring points:

  • Webhook delivery rate vs. expected change volume
  • Queue depth and processing lag
  • Gap between last webhook and last successful sync
  • Reconciliation job findings (should trend toward zero)
  • API throttling frequency

We alert if webhook delivery drops below expected baseline or if reconciliation consistently finds missed changes.

Recommendation: For your 30-second latency requirement, webhooks are the right choice but require proper infrastructure. The hybrid approach isn’t a compromise - it’s best practice for production systems. You get the latency benefits of webhooks with the reliability guarantees of polling. The added complexity is manageable and worthwhile for mission-critical data flows.