Real-time event processing vs batch ingestion: performance trade-offs for high-volume data

I want to discuss the performance trade-offs between real-time event processing and batch ingestion in AEC 2021 for high-volume customer interaction data. We’re currently processing 2-3 million events daily (web clicks, email opens, app interactions) and evaluating whether to stick with our 15-minute batch ingestion or move to real-time streaming.

Real-time streaming promises immediate data availability for personalization and triggers, but I’m concerned about throughput limits and infrastructure costs at our scale. Batch ingestion is efficient and reliable, but the 15-minute delay impacts our ability to respond to customer actions instantly. Hybrid processing models might offer the best of both worlds, but add complexity.

What have others experienced with real-time vs batch at high volumes? Where do you draw the line on which events go real-time vs batch? I’m interested in practical performance data and architecture insights.

At 2–3M events/day (~23–35 events/second sustained, with likely intraday peaks 3–5× that), you’re operating in a range where both approaches are technically viable in AEC, but the trade-offs sharpen considerably depending on your activation use cases.

Criteria Comparison

Criteria Batch (15-min) Real-Time Streaming Hybrid
Latency 15-min minimum lag Sub-minute (verify in your version) Tiered by event class
Throughput ceiling High — bulk processing absorbs spikes well Constrained by streaming ingestion limits; verify per-org caps in your contract Managed per tier
Infrastructure cost Lower compute, predictable Higher sustained cost; scales with event volume Moderate; complexity shifts cost to engineering
Data completeness Strong — retry logic and batch validation are mature Risk of out-of-order events, duplicate handling required Dependent on implementation
Personalization freshness Limited — stale profile state during session Enables in-session decisioning via Real-Time CDP profile updates Real-time for high-value signals; batch for enrichment
Operational complexity Low Medium-High — requires monitoring, dead-letter queues, schema validation Highest
Failure recovery Well-understood replay mechanisms More complex; failure mid-stream can cause profile inconsistency Mixed

Where Practitioners Typically Draw the Line

Real-time streaming is justified when the action window is shorter than your batch interval. Abandonment triggers, live chat escalation, and offer suppression during active sessions are cases where 15-minute lag destroys the use case entirely.

Batch ingestion remains appropriate for historical enrichment, aggregate scoring model inputs, offline campaign audience builds, and any signal where recency beyond 15 minutes doesn’t change the activation outcome.

Hybrid architecture (streaming for behavioral signals, batch for CRM/transactional data) is the dominant pattern at your scale. The typical implementation routes web click and app interaction events through the AEP Web SDK / Edge Network for real-time profile stitching, while email engagement and purchase history flow via scheduled Source Connectors or HTTP API batch uploads.

Practical Considerations at Your Volume

  • Verify your AEP streaming ingestion entitlements — throughput limits and event caps are contract-specific and often the first constraint teams hit when scaling.
  • At peak loads (marketing sends, campaign launches), batch provides natural backpressure absorption; streaming requires explicit capacity planning.
  • XDM schema validation failures at ingestion are more operationally painful in streaming than batch — get schema governance locked down before moving high-volume streams to real-time.
  • Profile fragmentation risk increases with streaming if identity resolution rules aren’t tuned — test merge policies against your identity graph before full rollout (verify behavior in your version).

Ultimately, this depends on context / your requirements — specifically, what your shortest viable action window is per event type, and whether your personalization use cases genuinely require sub-15-minute freshness or whether that’s an assumed requirement worth validating.


This draft is based on general Adobe Experience Cloud knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.

We moved from batch to real-time for critical events (purchases, cart abandonment, high-value page views) and kept batch for everything else. Real-time streaming in AEC 2021 handles about 5,000-10,000 events per second reliably, but beyond that you need careful tuning and scaling. Our total volume is similar to yours (2.5M daily) but only 15% goes through real-time stream, which keeps infrastructure manageable. The hybrid model works well - prioritize events that need immediate action and batch the rest.

The throughput vs latency trade-off is real. Real-time streaming gives you sub-second ingestion but at the cost of higher CPU usage, more complex error handling, and increased infrastructure costs. We measured 3-4x higher compute costs for real-time vs batch processing the same event volume. Batch ingestion is incredibly efficient - you can process millions of events with minimal resources because of optimized bulk operations. For most use cases, 15-minute latency is acceptable. The key question: what business value does real-time delivery actually provide for each event type?

We implemented a tiered hybrid model with three processing paths: ultra-fast (< 5 seconds) for transaction events and critical triggers, fast (1-2 minutes) for engagement events using micro-batches, and standard batch (15 minutes) for bulk activity data. This gives us 80% of the real-time benefit while keeping 70% of events in efficient batch processing. The micro-batch approach is a sweet spot - accumulate events for 60-90 seconds then process as a small batch. You get near-real-time latency with much better throughput than true streaming.

Don’t underestimate the operational complexity of real-time streaming. Batch jobs that fail can be easily retried - you just rerun the batch. Real-time stream failures are much harder to handle. You need dead letter queues, replay mechanisms, monitoring for lag, and strategies for handling backpressure. We’ve had incidents where a spike in event volume overwhelmed the real-time pipeline and caused hours of catch-up processing. Batch ingestion is much more forgiving of volume spikes and processing issues.

From an analytics perspective, batch ingestion actually has advantages. You get cleaner data because you can apply validation, deduplication, and enrichment across the entire batch before loading. Real-time streams often sacrifice data quality for speed - you’re writing events as they arrive with minimal transformation. We’ve seen 5-8% higher data quality scores with batch processing. For reporting and analytics use cases, the 15-minute delay is irrelevant and the improved data quality is valuable.

Consider your downstream consumers. If you move to real-time ingestion but your personalization engine or trigger rules only evaluate every 5 minutes anyway, you’ve gained nothing from the streaming infrastructure investment. We audited our use cases and found that only 20% actually benefited from sub-minute latency. The rest were perfectly fine with 15-minute batch windows. That analysis shaped our hybrid architecture and saved significant infrastructure costs.

Let me share comprehensive insights from our experience implementing hybrid processing at scale:

Real-Time Streaming Limits and Practical Throughput: AEC 2021’s real-time event processing has theoretical limits around 10,000-15,000 events per second per processing node, but practical sustainable throughput is lower - typically 5,000-8,000 events/sec to maintain reliability and leave headroom for spikes.

At your volume of 2-3 million events daily, that’s about 35-40 events per second average, which seems easily manageable. However, event distribution isn’t uniform. You likely see peaks of 200-500 events/sec during high-traffic periods (morning, promotional campaigns, etc.). Real-time streaming handles these bursts but requires careful capacity planning.

Key streaming limitations we’ve encountered:

  • Message ordering guarantees break down above 8K events/sec without partition key design
  • End-to-end latency increases from 1-2 seconds at normal load to 10-15 seconds under heavy load
  • Error rates spike above 0.5% when throughput exceeds 70% of capacity
  • Infrastructure costs scale nearly linearly with event volume (no economies of scale)

For your 2-3M daily volume, full real-time processing would cost 3-5x more than batch ingestion in infrastructure alone.

Batch Ingestion Efficiency and Performance: Batch ingestion is remarkably efficient because of bulk optimization opportunities:

  • Single database transaction for 50,000-100,000 events vs 50K individual transactions
  • Bulk validation and deduplication across the entire batch
  • Optimized compression and storage patterns
  • Simplified error handling and retry logic

Our batch processing handles 2,000-3,000 events per second per worker with minimal CPU usage (20-30%). The same infrastructure doing real-time streaming processes only 400-600 events/sec.

The 15-minute batch window is actually a sweet spot for balancing freshness and efficiency. Shorter windows (5 minutes) provide diminishing returns - you’re still not “real-time” but lose most efficiency gains. Longer windows (30+ minutes) create too much latency for modern customer expectations.

Batch ingestion excels at:

  • High-volume, low-urgency events (page views, email opens, routine activity)
  • Events requiring complex transformation or enrichment
  • Historical data loads and backfills
  • Scenarios where exactly-once processing is critical

Hybrid Processing Models - Architecture Patterns: The optimal architecture uses event classification to route to appropriate processing paths:

Tier 1 - Real-Time Stream (< 5 seconds latency):

  • Transaction events (purchases, subscriptions, cancellations) - immediate business impact
  • Cart abandonment triggers - time-sensitive recovery opportunity
  • High-value customer actions - VIP segment interactions requiring immediate response
  • Critical error events - system alerts, payment failures Volume target: 5-10% of total events

Tier 2 - Micro-Batch (1-3 minutes latency):

  • Engagement events (content views, downloads, form submissions)
  • Email clicks and opens - fast enough for follow-up campaigns
  • Product searches and browsing patterns
  • Social media interactions Volume target: 15-20% of total events

Implementation: Accumulate events in memory buffer for 60-90 seconds, then process as mini-batch

Tier 3 - Standard Batch (15 minutes latency):

  • Bulk activity data (page views, session tracking, analytics events)
  • Non-urgent engagement metrics
  • Historical data synchronization
  • Reporting and analytics feeds Volume target: 70-75% of total events

Implement event classification rules in your ingestion API:

  • Event type + customer segment determines processing tier
  • High-value customers get more real-time processing
  • Same event type (e.g., page view) routes differently based on page importance

The micro-batch tier is the secret weapon - it provides 90% of real-time benefits at 40% of the infrastructure cost. Events accumulate for 60-90 seconds, then process as a small batch with bulk operations. Latency is low enough for most personalization needs while maintaining batch efficiency.

Performance Data and ROI Analysis: Based on our implementation with similar volume (2.8M events/day):

Full Real-Time Architecture:

  • Average latency: 2-3 seconds
  • Infrastructure cost: $8,500/month
  • Operational complexity: High (24/7 monitoring required)
  • Data quality score: 92%

Full Batch Architecture:

  • Average latency: 12-15 minutes
  • Infrastructure cost: $1,800/month
  • Operational complexity: Low (standard job monitoring)
  • Data quality score: 97%

Hybrid Architecture (10% real-time, 20% micro-batch, 70% batch):

  • Average latency: 6-8 minutes (weighted across tiers)
  • Infrastructure cost: $3,200/month
  • Operational complexity: Medium
  • Data quality score: 96%
  • Business value: 85% of full real-time benefit at 38% of cost

The hybrid model delivers optimal ROI - you get real-time processing where it matters most while maintaining efficiency for bulk data.

Practical Implementation Recommendations: Start with batch ingestion for everything and incrementally move high-value events to real-time. This de-risks the migration and lets you measure actual business impact.

Implement event prioritization scoring: assign each event type a business value score (1-10) and urgency score (1-10). Events with combined score > 15 go real-time, 10-15 go micro-batch, < 10 stay batch.

Build comprehensive monitoring for both paths. Track: throughput, latency, error rates, data quality metrics, infrastructure costs. This data drives ongoing optimization of which events belong in which tier.

Design for degradation. If real-time stream falls behind or fails, automatically route events to batch processing. Better to have 15-minute latency than lost data.

The key insight: real-time processing is a premium service. Apply it selectively to events that truly benefit from immediate processing. For your 2-3M daily volume, targeting 10-15% real-time processing gives you the responsiveness customers expect while keeping infrastructure costs reasonable and operations manageable.