We’re running about 400K contacts and 120K accounts in our Salesforce org, and duplicates are becoming a real problem. Our current setup uses basic duplicate rules—exact match on email, fuzzy on name and company—but we’re missing a lot of subtle variations. Misspellings, abbreviations, different formats for the same company name (“Tech Corp” vs “Technology Corporation Inc”), you know the drill.
We’re also struggling with account hierarchies. Parent-child relationships don’t update when companies get acquired or restructured, and our account teams spend hours each week researching which subsidiary belongs where. We’ve looked at a few third-party dedup tools, and some are purely rule-based while others use machine learning. The ML vendors talk about continuous improvement and handling complexity, but they also require training data and some black-box logic that makes our compliance team nervous.
I’m curious what others have done when scaling past the basic duplicate rules. Did you stick with rule-based matching and just invest more time building better rules? Or did you go the ML route, and if so, how did you handle the training phase and explainability requirements? Also interested in any strategies for keeping account hierarchies current—manual processes, third-party enrichment feeds, or something more automated.
At 400K contacts / 120K accounts, you’ve hit the ceiling where native Duplicate Rules + Matching Rules stop being sufficient. Here’s a breakdown of where each approach lands and what’s working at scale.
Native Rule-Based (Salesforce Matching Rules)
The built-in engine supports fuzzy matching via Edit Distance and Acronym match methods, but abbreviation normalization (“Tech Corp” vs “Technology Corporation Inc”) requires pre-processing or custom logic you have to maintain yourself. Complexity compounds fast—each edge case becomes a new rule, and rule sprawl creates conflicts. Fine for orgs under ~100K records with clean input; beyond that, maintenance cost exceeds value.
ML-Based Matching: What It Actually Requires
The ML vendors aren’t wrong about handling variation better, but the prerequisites matter:
Training data: You need a labeled golden dataset—confirmed match/non-match pairs. Minimum viable is typically 1,000–5,000 labeled pairs; more if your data is noisy. Pulling this from your existing records (using your current rules as a bootstrap, then manually reviewing edge cases) is the standard approach.
Explainability: Most mature vendors (verify in your version) expose field-level confidence scores per match—Email match: 95%, Name match: 72%, Address match: 60%—which satisfies most compliance review requirements. “Black box” is less accurate than “auditable at field level.” Ask vendors specifically for a match rationale output in their demo.
Continuous learning: This only kicks in if you feed back merge/unmerge decisions as signals. Without a closed feedback loop, the model doesn’t improve.
Account Hierarchies at Scale
Manual maintenance is unsustainable. The realistic options:
Third-party enrichment: D&B Hoovers, Cognism, or ZoomInfo connectors can push D-U-N-S-linked hierarchy data directly into the ParentId field on Account. Acquisitions and restructures propagate via scheduled sync jobs. Verify data freshness SLAs before committing—some feeds lag 30–90 days.
Ownership field standardization: Before any enrichment layer, normalize your Account.Name and add a D-U-N-S Number or External ID field as the stable anchor. Matching on a mutable name field is the root cause of hierarchy drift.
Change event triggers: Use Change Data Capture on Account to fire hierarchy validation logic when ParentId updates, rather than batch jobs.
Practical path: Bootstrap ML training data from your existing rule matches, use a vendor that exposes field-level confidence, and anchor hierarchies to an immutable external identifier before layering enrichment. The compliance concern is addressable—it’s a vendor selection and demo-script question, not a fundamental blocker.
This draft is based on general salesforce knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.
We went the ML route about 18 months ago with a vendor that uses similarity scoring rather than hard rules. The training phase wasn’t too bad—we labeled maybe 25–30 record pairs as match/no-match, and the model picked up patterns pretty quickly. The big win for us was adaptability. When we merged with another company and suddenly had a whole new set of naming conventions and data formats, the ML model adjusted without us rewriting a hundred rules. Explainability was less of an issue than we expected because the vendor provides confidence scores and lets us set thresholds for auto-merge vs. manual review.
For account hierarchies, we’ve had good results with a combination of third-party enrichment and manual review workflows. We use an API that pulls corporate structure data from a reference provider, then surface recommendations to account owners through a custom Lightning component. They can accept or reject each suggestion, and accepted changes flow back into Salesforce through a scheduled job. It’s not fully automated, but it cut research time by more than half and keeps hierarchies much more current than our old manual approach.
One thing we learned the hard way: even ML-based dedup depends on decent baseline data quality. We tried to roll out a machine learning tool when our data was a mess—missing fields, inconsistent formatting, no validation rules—and the model couldn’t learn reliable patterns. We had to step back, implement field-level validation, standardize addresses and phone numbers, and clean up a few thousand obvious duplicates manually before the ML approach started delivering value. The bootstrap problem is real.
Have you considered a hybrid approach? We use rule-based matching for high-confidence scenarios—exact email match, phone number formatting variants, that sort of thing—and ML for the ambiguous cases. The rule-based engine handles maybe 60% of duplicates automatically, and the ML model picks up the rest. This way we get fast, transparent decisions where the logic is obvious, and the ML model focuses on the edge cases where it actually adds value. Also keeps computational costs lower than running everything through inference.
From a compliance perspective, the key question is auditability. If you’re in a regulated industry or dealing with GDPR/CCPA, you need to be able to explain why records were merged or why a hierarchy change was made. Some ML vendors provide chain-of-thought explanations or reasoning logs that satisfy auditors, but not all do. Make sure whatever tool you pick can produce a paper trail that shows the data sources, similarity scores, and decision logic. We rejected one vendor because their model was too much of a black box and our legal team wasn’t comfortable with it.
For account hierarchies at scale, vector embeddings and similarity search can be a game changer. We embed account names using a sentence transformer model, store them in a vector database, and run similarity queries to find potential parent matches. Combine that with web search results and a third-party API for corporate structure data, then feed everything into an LLM to generate recommendations with reasoning. It’s not fully automated—account managers review and approve changes—but it’s way faster than manual research and handles spelling variations and rebranded entities that exact-match lookups miss.
One practical tip: whichever approach you choose, set up closed-loop feedback. Sales reps and account managers will spot bad merges or incorrect hierarchy assignments faster than anyone else. Build a simple workflow where they can flag issues, and feed that back into your rules or training data. We’ve improved our match accuracy by about 15% just from incorporating user corrections over six months. The model or rule set gets smarter, and the team feels like they have a voice in how the system works.