Two category labels look identical on screen, yet a program treats them as different strings. Before deleting one as a duplicate, inspect how the text is represented and decide what equivalence means for the dataset.
The Python unicodedata documentation describes Unicode normalization forms. A character such as an accented letter can have a composed representation or a sequence containing a base letter and a combining mark.
Use a synthetic example before touching real records. In Python, compare "\u00e9" with "e\u0301", then compare their NFC-normalized forms. The visible accent is not enough to explain the raw string comparison; the code points matter.
Keep the original while testing a matching key
Create a separate normalized field and count how many previously distinct values converge. Review those collisions. Normalization is not the same as lowercasing, removing accents or translating names, and those additional changes require their own justification.
Avoid applying compatibility normalization indiscriminately to identifiers where distinctions may be meaningful. Choose the form according to the data contract, not merely because it produces fewer categories.
Record the transformation and the software used. A useful quality check includes examples that should match and examples that must remain distinct. The objective is a documented comparison rule, while retaining enough original information to investigate an unexpected match.
Image: an editorial illustration, not a documentary photograph.

