Here’s a comprehensive solution that addresses all three focus areas without requiring a full multi-region architecture:
First, addressing why jobs fail only during AWS issues while manual updates succeed: The root cause is connection timeout sensitivity. Manual stock updates are short-lived transactions (typically under 5 seconds) that complete before any network disruption impacts them. Your batch jobs, however, process hundreds or thousands of stock records in a single session, holding database connections for 10-30 minutes. When AWS experiences regional connectivity issues - even brief 30-60 second disruptions - these long-held connections exceed TCP timeout thresholds and get terminated by the network layer. CloudSuite’s batch framework doesn’t always detect these connection drops gracefully, causing the job to fail without proper error handling.
Second, implementing automatic recovery since manual updates succeed: Configure intelligent retry logic at the batch scheduler level. In CloudSuite, navigate to System Administration > Batch Scheduler > Job Definitions. Edit your stock update job and enable these settings:
- Retry on Failure: Enabled
- Maximum Retry Attempts: 3
- Retry Delay Strategy: Exponential Backoff (5min, 15min, 30min)
- Retry Conditions: Network errors, connection timeouts, database unavailable
- Checkpoint Interval: 100 records
The checkpoint setting is critical - it tells the job to commit progress every 100 stock records. If the job fails at record 450, the retry will resume from record 400 instead of starting over. This makes retries much more likely to succeed because they have less work to complete before the next potential disruption.
Additionally, implement connection health checks within the job itself. Add a pre-execution validation step that tests database connectivity and AWS service availability. If the health check fails, the job should delay execution by 10 minutes and retry the health check rather than attempting to process records with degraded connectivity.
Third, solving the alerting gap for failed jobs: The lack of notifications is dangerous because inventory mismatches compound over time. Implement a two-tier alerting strategy:
Immediate Alerts: Configure ION workflow notifications for real-time job failures. In ION Desk, create a workflow document flow that subscribes to the ‘BatchJobFailed’ event. Route this to your operations team via email and SMS. The alert should include:
- Job name and scheduled time
- Failure reason and error code
- Number of records processed before failure
- Automatic retry status
Daily Summary: Create a scheduled report that runs every morning at 7 AM listing all batch jobs from the previous 24 hours with their completion status. This catches any failures that might have been missed by the immediate alerts and provides a comprehensive view of batch processing health.
Implement AWS CloudWatch integration for infrastructure-level monitoring. Set up a CloudWatch alarm that monitors your CloudSuite database connection pool metrics. When connection failures spike above 5% within a 5-minute window, trigger a preventive alert that warns your team about potential AWS issues before batch jobs fail. This gives you time to manually delay job execution or switch to manual processing if needed.
With these configurations in place, your batch jobs will automatically retry after transient AWS failures, resume from checkpoints rather than restarting completely, and notify your team immediately if all retries are exhausted. This eliminates the multi-day discovery delays you’re currently experiencing and brings batch job reliability in line with your manual update success rate.
This draft is based on general Infor CloudSuite knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.