Optimized loyalty points batch processing using Bulk API

I want to share our implementation story of optimizing loyalty points processing for a retail client with 2 million active loyalty members. They were processing points accrual and redemption using individual API calls, which was taking 8-12 hours nightly and frequently failing due to timeouts.

We redesigned the entire architecture using Salesforce Bulk API with parallel processing and asynchronous Apex monitoring. The results were dramatic: processing time dropped from 8-12 hours to under 90 minutes, failure rate went from 15% to under 1%, and we freed up significant system resources for other operations. The rewards processing became reliable enough to run multiple times per day instead of just overnight.

The key architectural decisions were using Bulk API 2.0 for data loading, implementing proper batch size tuning based on transaction complexity, and building a custom monitoring framework using asynchronous Apex to track job progress and handle failures gracefully. Happy to share technical details of the implementation if others are tackling similar high-volume batch processing challenges.

What monitoring framework did you build? We use standard Bulk API job monitoring but it doesn’t give us enough visibility into where failures happen or which batches are slow. Would love to hear about your asynchronous Apex approach for tracking job progress. Also, did you implement any retry logic for failed batches, or do you reprocess the entire job?

This is exactly the kind of optimization story we need more of. 8 hours down to 90 minutes is incredible. I’m curious about your batch size tuning approach - did you use fixed batch sizes or dynamic sizing based on transaction types? We’re processing about 500K loyalty transactions daily and struggling with the right batch size. Too large and we hit governor limits, too small and the overhead kills performance.

The failure rate improvement is just as impressive as the speed gain. 15% to under 1% suggests you fixed some fundamental issues beyond just using Bulk API. Were the original failures due to timeout issues, data validation errors, or governor limits? And how did your architecture changes address each failure mode? Understanding the failure patterns would help us diagnose our own batch processing issues.

Running batch processes multiple times per day instead of just overnight is a huge business value add. Real-time or near-real-time points visibility improves customer satisfaction significantly. What was the trigger for running jobs multiple times daily - was it time-based scheduling, or did you implement event-driven processing where jobs run when transaction volumes reach certain thresholds? We want to move away from overnight-only processing but need to balance system load.

Thanks for the interest! I’ll break down the technical implementation in detail:

Bulk API Parallel Processing Architecture:

We migrated from Bulk API 1.0 to Bulk API 2.0 specifically for the parallel processing capabilities. Bulk API 2.0 processes batches in parallel automatically, whereas 1.0 was sequential. This alone cut processing time by 60%.

Implementation pattern:

// Pseudocode - Bulk API 2.0 job creation:
1. Prepare CSV data file with loyalty transactions
2. Create job using Bulk API 2.0 endpoint: POST /services/data/v58.0/jobs/ingest
3. Set job parameters: object=LoyaltyTransaction__c, operation=insert, lineEnding=LF
4. Upload CSV data in chunks (max 100MB per upload)
5. Close job to begin processing
// Jobs process in parallel across multiple batches automatically

The parallel processing means multiple batches execute simultaneously rather than waiting for each to complete. For 2 million records, we typically see 15-20 batches running in parallel, dramatically reducing total processing time.

Batch Size Tuning Strategy:

This was critical for optimization. We implemented dynamic batch sizing based on transaction complexity:

Simple transactions (points accrual only): 5,000 records per batch

Complex transactions (redemptions with multiple line items): 2,000 records per batch

Mixed transactions: 3,000 records per batch

The tuning process involved:

  1. Starting with Salesforce’s recommended 10,000 records per batch
  2. Monitoring batch execution times and failure rates
  3. Reducing batch size when we saw governor limit exceptions (CPU time, heap size)
  4. Testing different sizes to find the optimal balance between throughput and reliability

For your 500K daily transactions, I’d recommend starting with 3,000 per batch and adjusting based on your specific transaction complexity. Monitor the Bulk API job details to see which batches are slow or failing.

Asynchronous Apex Monitoring Framework:

Standard Bulk API monitoring was insufficient for our needs. We built a custom monitoring system using scheduled Apex:

// Pseudocode - Monitoring job implementation:
1. Scheduled Apex runs every 5 minutes during processing windows
2. Query Bulk API jobs: GET /services/data/v58.0/jobs/ingest/{jobId}
3. Parse job status: numberRecordsProcessed, numberRecordsFailed
4. Store metrics in custom object: BulkJobMonitor__c
5. Send alerts if failure rate exceeds 2% or processing stalls
// Detailed logging enables root cause analysis

Key monitoring metrics we track:

  • Total records processed vs expected
  • Processing rate (records per minute)
  • Failure rate by batch
  • CPU time consumption per batch
  • Heap size usage patterns

This monitoring identified bottlenecks we couldn’t see with standard tools. For example, we discovered certain transaction types consumed 3x more CPU time, leading us to process them in smaller batches.

Failure Handling and Retry Logic:

We implemented intelligent retry logic rather than reprocessing entire jobs:

  1. After job completion, query failed records using Bulk API results endpoint
  2. Categorize failures: validation errors vs governor limit exceptions vs data integrity issues
  3. Validation errors → log for manual review (usually data quality issues)
  4. Governor limit exceptions → retry in smaller batches (1,000 records instead of 5,000)
  5. Data integrity issues → pause processing, alert support team

Failed batches retry automatically up to 3 times with exponential backoff. This reduced our failure rate from 15% to under 1% because transient issues (temporary governor limit hits, brief system slowdowns) resolve on retry without manual intervention.

Dependency Management with Parallel Processing:

Transaction dependencies were tricky with parallel processing. Our solution:

Sequenced job stages:

  • Stage 1: Process all points accruals (can run fully in parallel)
  • Stage 2: Wait for Stage 1 completion, then process redemptions (parallel within stage)
  • Stage 3: Process adjustments and corrections (parallel within stage)

We use platform events to coordinate between stages:

// Pseudocode - Stage coordination:
1. Stage 1 Bulk API job completes
2. Monitoring Apex publishes platform event: AccrualProcessingComplete__e
3. Event-triggered flow starts Stage 2 Bulk API job
4. Stages remain parallel internally but sequential across stages

This maintains parallelism benefits while respecting business logic dependencies. Each stage processes thousands of records in parallel, but stages execute sequentially to ensure data integrity.

Multi-Daily Processing Implementation:

We moved from overnight-only to 4x daily processing:

  • 6 AM: Overnight transaction batch (largest volume)
  • 12 PM: Morning transaction batch
  • 6 PM: Afternoon transaction batch
  • 11 PM: Evening transaction batch

Trigger mechanism is hybrid:

  • Time-based: Scheduled jobs at fixed times
  • Volume-based: If transaction queue exceeds 50,000 records, trigger immediate processing
  • Event-based: High-priority transactions (VIP customer redemptions) trigger dedicated small batch jobs

This approach balances system load across the day while providing near-real-time processing for most transactions. Customers see points updates within 6 hours maximum, usually within 2-3 hours.

Performance Results:

Before optimization:

  • Processing time: 8-12 hours (single overnight batch)
  • Failure rate: 15% (300,000 records failed daily)
  • System resource usage: 80%+ during processing
  • Customer points visibility: Next day only

After optimization:

  • Processing time: 90 minutes per batch (4x daily)
  • Failure rate: <1% (under 20,000 records failed daily)
  • System resource usage: 40-50% during processing (freed resources for other operations)
  • Customer points visibility: 2-6 hours typical

The business impact was significant - customer satisfaction scores for loyalty program increased by 23% due to faster points visibility, and the support team’s workload decreased by 40% because fewer customers called about missing points.

Key Takeaways:

  1. Bulk API 2.0 parallel processing is essential for high-volume operations
  2. Dynamic batch sizing based on transaction complexity optimizes throughput
  3. Custom monitoring provides visibility that standard tools lack
  4. Intelligent retry logic handles transient failures automatically
  5. Staged processing balances parallelism with business logic dependencies

For anyone implementing similar high-volume batch processing, start with Bulk API 2.0, implement comprehensive monitoring, and tune batch sizes based on your specific data patterns. The investment in custom monitoring and retry logic pays off quickly in reliability and reduced manual intervention.

Bulk API 2.0 is definitely the right choice for this volume. The parallel processing capabilities are game-changing compared to the original Bulk API. One challenge we faced was handling dependencies between different types of transactions - points accrual needs to complete before redemptions can process. How did you sequence your batch jobs to handle these dependencies while maintaining parallel processing efficiency?