One-Hot Encoding Turns Categories into Separate Yes-or-No Columns

Conceptual desk scene with a notebook, a pressed leaf and blank cards beside a laptop.
Conceptual illustration created with AI; not a research result or laboratory photograph.

Assigning red, green, and blue the numbers one, two, and three can accidentally suggest an order that the categories do not possess. One-hot encoding offers another representation: a separate indicator column for each category.

The scikit-learn OneHotEncoder documentation describes this transformation and the options for handling categories. The numerical representation should match the modelling question, rather than invent a ranking merely because numbers are convenient.

Work through three fictional records

Choose the column order red, green, blue. A red record becomes 1, 0, 0; a blue record becomes 0, 0, 1. Keep the column names beside the values. Without that mapping, the sequence alone does not identify the colour.

Then imagine receiving yellow after the encoder was fitted to the original categories. Decide deliberately how unknown categories are handled using the documented configuration. Do not silently assign a familiar category just to make the input fit.

Fit preprocessing rules using the appropriate training data when building a predictive workflow, and apply the established mapping consistently to later records. A different column order can change the meaning of an otherwise identical row.

Finally, check whether the field is truly unordered. A ranked satisfaction scale asks a different representation question from a set of colour names. Encoding is a modelling choice with assumptions, not a universal cleanup step.

Editorial illustration from this site’s image library; not documentary evidence of the example or object discussed.