Add this validation script to your ETL preprocessing pipeline before the ECN Loader runs:
import pandas as pd
import re
def normalize_bom_items(df):
df['item_id'] = df['item_id'].str.upper().str.strip()
df['item_id'] = df['item_id'].str.replace(r'\s+', ' ', regex=True)
return df
def validate_duplicates(df):
dupes = df.groupby(['ecn_id', 'item_id']).size()
return dupes[dupes > 1]
This addresses all three critical focus areas:
BOM Deduplication: The script identifies true duplicates after normalization. Run this validation before loader execution to catch issues early. If duplicates remain after normalization, they’re legitimate data quality problems requiring business decision on which record to keep.
Loader Mapping Validation: Update your ECN Loader mapping configuration to apply the same normalization rules. In the loader XML mapping file, add transformation functions:
<Mapping source="ItemID" target="item_number">
<Transform function="uppercase"/>
<Transform function="trim"/>
<Transform function="normalize_whitespace"/>
</Mapping>
This ensures the loader applies identical normalization during import, preventing mismatches between your preprocessing and the actual load operation.
Whitespace/Case Normalization: The uppercase and trim operations handle the specific example you showed (‘PART-12345’ vs ‘part-12345’). The regex whitespace normalization catches hidden issues like tabs, multiple spaces, or non-breaking spaces that visual inspection misses.
Additional considerations:
-
Occurrence vs Item Duplicates: Verify your loader configuration distinguishes BOM occurrences (same part appearing multiple times with different quantities/positions) from true duplicates. Set the ‘allowMultipleOccurrences’ flag to true in your ECN Loader config.
-
Attribute Preservation: Ensure your mapping preserves differentiating attributes like find numbers, reference designators, and position codes. These make occurrences unique even when they reference the same part.
-
Validation Reporting: Generate a pre-migration report showing all normalized duplicates with their original values. This helps the business team review and resolve ambiguous cases before migration.
-
Incremental Testing: Test with a small batch (100-200 ECNs) first to validate the normalization rules don’t inadvertently merge legitimately distinct items.
For your 15,000 ECN migration, implement this as a three-stage process: normalize source data, validate for duplicates, then load. This approach reduced our migration errors from 23% to under 2% and made the remaining issues easy to identify and resolve manually.
This draft is based on general Teamcenter knowledge. It has not been verified against your specific version and environment. Practitioners: verify the steps and share your experience below.