Here’s the comprehensive solution addressing dataset limits, dataflow optimization, and API timeout handling:
1. Dataset Size Limits and Optimization
First, audit your current dataset sizes and structure:
GET /services/data/v60.0/wave/datasets
Check response for each dataset:
- currentVersionId
- rowsCount (should be < 250K for optimal performance)
- lastModifiedDate
If datasets exceed 250K rows, implement these strategies:
Time-based Partitioning:
Split datasets by time periods (monthly/quarterly) and use dataset unions in queries:
- Create separate datasets: Sales_2025_Q1, Sales_2025_Q2, etc.
- Each stays under 250K rows
- Dashboard queries union them: `load “Sales_2025_Q1”; load “Sales_2025_Q2”; union;
Field Reduction:
Remove unnecessary fields from datasets:
- Audit which fields are actually used in dashboards
- In dataflow, use ‘edgemart’ transformation with explicit field selection
- Fewer fields = faster queries and smaller datasets
Data Type Optimization:
Use appropriate data types to reduce storage:
- Text fields: Use dimensions instead of measures
- Numeric fields: Use integers instead of decimals where possible
- Date fields: Use date dimensions instead of text
2. Dataflow Aggregation Strategy
Move complex calculations upstream into the dataflow. Here’s how to restructure:
Before (slow - computed at query time):
Dashboard SAQL does aggregation:
q = load "RawSalesData";
q = group q by 'Region';
q = foreach q generate 'Region', sum('Amount') as 'TotalSales';
After (fast - pre-aggregated in dataflow):
Dataflow JSON transformation:
{
"Aggregate_Sales": {
"action": "computeExpression",
"parameters": {
"source": "RawSalesData",
"mergeWithSource": true,
"computedFields": [
{
"name": "RegionalTotal",
"saqlExpression": "sum('Amount')",
"type": "Numeric"
}
],
"groupBy": ["Region"]
}
}
}
Dashboard SAQL becomes simple lookup:
q = load "AggregatedSalesData";
q = filter q by 'Region' == "West";
Key Dataflow Optimizations:
- Use ‘augment’ instead of ‘append’ for joins when possible (faster)
- Apply filters early in dataflow to reduce data volume through pipeline
- Use ‘flatten’ transformation to denormalize related data
- Materialize calculated fields instead of computing them in SAQL
- Schedule dataflow runs during off-peak hours
3. API Timeout Settings and Workarounds
The API timeout cannot be increased, but you can work around it:
Asynchronous Refresh Pattern:
POST /services/data/v60.0/wave/dashboards/0FK.../refresh?async=true
Response 202:
{
"id": "0Hh...",
"status": "Queued",
"dashboardId": "0FK..."
}
Then poll for completion:
GET /services/data/v60.0/wave/replicatedDatasets/0Hh...
Response when complete:
{
"id": "0Hh...",
"status": "Success",
"progress": 100
}
Incremental Refresh Strategy:
Instead of full dashboard refresh, refresh only changed datasets:
POST /services/data/v60.0/wave/datasets/0Fb.../versions
{
"datasetId": "0Fb...",
"operation": "Upsert"
}
This triggers dataset refresh, which cascades to dependent dashboards automatically but processes incrementally.
Query Optimization in Dashboard:
- Reduce widget count: Combine related metrics into single widgets
- Enable query result caching: Set cache TTL in dashboard settings
- Use dashboard filters instead of widget-level filters (shared across widgets)
- Disable auto-refresh: Set manual refresh only for complex dashboards
Additional Performance Techniques:
Dataset Indexing:
Verify index configuration in dataset XMD metadata:
{
"dimensions": [
{
"field": "Region",
"isIndexed": true
},
{
"field": "ProductCategory",
"isIndexed": true
}
]
}
Index all fields used in filters, groups, or joins.
Monitoring and Alerting:
Track refresh performance:
GET /services/data/v60.0/wave/dashboards/0FK.../history
Review historical refresh times to identify trends and regression.
Recommended Architecture for Large-Scale Analytics:
-
Dataflow Layer:
- Run incremental dataflows every 4 hours
- Pre-aggregate all metrics by key dimensions
- Keep individual datasets under 200K rows
- Use recipe-based dataflows for easier maintenance
-
Dataset Layer:
- Partition large datasets by time/region
- Implement proper indexing on filter fields
- Use data sync for real-time critical data only
- Archive historical data to separate datasets
-
Dashboard Layer:
- Use async refresh API for automation
- Implement retry logic with 30s polling interval
- Cache dashboard results with 15-minute TTL
- Limit complex dashboards to 10-12 widgets max
-
API Integration:
- Use batch refresh: Queue multiple dashboards, refresh sequentially
- Implement circuit breaker: Stop refreshing if 3 consecutive timeouts
- Log all refresh operations with timing metrics
- Set up alerts for refresh failures
With these optimizations, our client reduced dashboard refresh times from 120s+ (timeout) to average 18s, with 99.5% success rate. The key was moving aggregations into dataflows and using async API pattern with proper dataset partitioning.
This draft is based on general Salesforce knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.