Let me provide a comprehensive solution covering all four critical areas for reliable webhook delivery:
Webhook Retry Logic:
Implement exponential backoff with jitter in your webhook receiver. Event Grid already handles retries on its side, but you need defensive coding in your endpoint:
app.post('/webhook/registration', async (req, res) => {
// Return 200 immediately to acknowledge receipt
res.status(200).send();
// Process asynchronously with retry logic
await processWithRetry(req.body, {
maxAttempts: 5,
baseDelay: 1000,
maxDelay: 30000
});
});
Key retry configuration:
- Event Grid retry policy: Set max delivery attempts to 30 with event TTL of 1440 minutes (24 hours)
- Implement circuit breaker pattern: After 5 consecutive failures, pause webhook processing for 5 minutes
- Use exponential backoff: delay = min(maxDelay, baseDelay * 2^attempt)
- Add jitter to prevent thundering herd: delay += random(0, 1000)
Azure Event Grid Integration:
Optimize your Event Grid subscription configuration:
{
"destination": {
"endpointType": "webhook",
"properties": {
"endpointUrl": "https://your-api.com/webhook/registration",
"maxEventsPerBatch": 1,
"preferredBatchSizeInKilobytes": 64
}
},
"retryPolicy": {
"maxDeliveryAttempts": 30,
"eventTimeToLiveInMinutes": 1440
},
"deadLetterDestination": {
"endpointType": "storageBlob",
"properties": {
"resourceId": "/subscriptions/{sub}/resourceGroups/{rg}/providers/Microsoft.Storage/storageAccounts/{account}",
"blobContainerName": "event-deadletter"
}
}
}
Enable diagnostic logging to track delivery metrics:
- DeliveryFailures
- PublishSuccesses
- DeadLetteredEvents
- MatchedEvents
Set up alerts when DeliveryFailures exceed 5% of total events.
Dead-Letter Queue Implementation:
Implement comprehensive dead-letter handling:
- Configure Azure Storage dead-letter destination (shown above)
- Create an Azure Function triggered by new blobs in the dead-letter container
- Implement intelligent replay logic:
// Pseudocode for dead-letter processor:
1. Read failed event from blob storage
2. Check failure reason and count
3. If transient error (timeout, 5xx): retry immediately
4. If persistent error (4xx): log for manual review
5. If retry count > threshold: send to admin queue
6. Track retry attempts in event metadata
- Set up monitoring dashboard showing:
- Dead-letter queue depth
- Event age in dead-letter
- Replay success rate
- Common failure patterns
Idempotency Handling:
Implement robust duplicate detection:
const processedEvents = new Map(); // In production, use Redis
async function processEvent(event) {
const eventId = event.id;
// Check if already processed
if (await isProcessed(eventId)) {
console.log(`Skipping duplicate event: ${eventId}`);
return;
}
// Mark as processing (with TTL)
await markProcessing(eventId, ttl: 86400); // 24 hours
try {
// Your business logic here
await updateRegistration(event.data);
// Mark as completed
await markCompleted(eventId);
} catch (error) {
// Remove processing lock to allow retry
await clearProcessing(eventId);
throw error;
}
}
Use Redis or Azure Cache for distributed idempotency tracking:
- Store event ID as key with processed timestamp as value
- Set TTL to 25 hours (longer than Event Grid’s 24-hour retry window)
- Use atomic operations (SET NX) to prevent race conditions
Additional Best Practices:
- Webhook Validation: Implement Event Grid webhook validation:
if (req.body[0].eventType === 'Microsoft.EventGrid.SubscriptionValidationEvent') {
return res.json({ validationResponse: req.body[0].data.validationCode });
}
-
Security:
- Validate Event Grid signature headers
- Use managed identities for authentication
- Implement IP allowlisting for webhook endpoint
-
Monitoring:
- Track end-to-end latency from event creation to processing completion
- Monitor webhook endpoint availability (99.9% target)
- Alert on dead-letter queue depth > 100 events
- Create dashboard showing registration update success rate
-
Testing:
- Implement chaos engineering: randomly fail webhook deliveries to test retry logic
- Load test with 1000 concurrent registration updates
- Verify idempotency by sending duplicate events intentionally
With this architecture, you’ll achieve 99.99% delivery reliability with sub-second latency for successful deliveries and automatic recovery from transient failures. The dead-letter queue ensures no events are lost, and idempotency handling prevents duplicate processing even during retry storms.
This draft is based on general Microsoft Dynamics 365 Sales knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.