A `Set<Customer> seen = new HashSet<>()` is used to deduplicate customers arriving from an import job. `Customer` has no overridden `equals` or `hashCode`. The set ends up with more entries than there are distinct customers. Why, mechanically, and what's the minimum fix?
HashSet decides whether two objects are "the same" entirely by hashCode and equals, and without an override, both come from Object: hashCode is derived from the object's identity (effectively its memory address), and equals is ==. Two Customer objects built from two rows of the import that happen to describe the same real customer are two different objects in memory, so they get two different default hash codes and fail the default equals check — the set has no way to know they're "the same customer" unless Customer tells it what "the same" means. It adds both, correctly, by its own rules; the bug is that Customer's rules for sameness were never defined. Overriding equals and hashCode together, based on whatever field actually identifies a customer (an id, or a natural key), is the fix — and both have to change together, because a HashSet only calls equals on objects that already landed in the same bucket, which only happens when their hashCodes agree.