Removing duplicates from a list in Python is one of the most common data-cleaning tasks developers face when working with real-world data. But whether you are processing user input, preparing data for analysis, or building a clean dataset for machine learning, duplicate values can cause incorrect counts, skewed results, and inefficient memory usage. This article explains practical ways to remove duplicates from a list in Python, including methods that preserve order, handle unhashable items, and improve performance No workaround needed..
Introduction
In Python, a list is an ordered, mutable collection that can contain repeated values. For example:
data = [10, 20, 10, 30, 20, 40]
The list above contains duplicate values: 10 and 20 appear more than once. If your program needs unique values only, you need a reliable way to remove
Removing Duplicates with a Set
The simplest approach to eliminate duplicates is to convert the list to a set, which inherently contains only unique elements. That said, this method does not preserve the original order of elements:
data = [10, 20, 10, 30, 20, 40]
unique_data = list(set(data))
print(unique_data) # Output: [10, 20, 30, 40] or any order
While effective for small datasets, this method is unsuitable when order matters. For ordered deduplication, alternative strategies are required And it works..
Preserving Order with a Loop
To maintain the original order while removing duplicates, iterate through the list and track seen elements with a helper set:
def remove_duplicates_ordered(input_list):
seen = set()
result = []
for item in input_list:
if item not in seen:
seen.add(item)
result.append(item)
return result
data = [10, 20, 10, 30, 20, 40]
print(remove_duplicates_ordered(data)) # Output: [10, 20, 30, 40]
This approach ensures uniqueness while preserving insertion order, making it ideal for scenarios like cleaning user
…cleaning user input, maintaining audit trails, or preparing time‑series data where the chronological order of events must stay intact.
Handling Unhashable Items
When the list contains mutable objects such as dictionaries, lists, or sets, the simple seen = set() trick fails because these types cannot be hashed. A common workaround is to replace the hash‑based seen with a list (or another suitable container) and rely on equality checks:
def remove_duplicates_unhashable(seq):
seen = []
result = []
for item in seq:
if item not in seen: # linear search; acceptable for small‑to‑moderate sizes
seen.append(item)
result.append(item)
return result
If the unhashable elements are themselves hashable after a transformation (e.g., converting inner lists to tuples), you can keep the fast set lookup:
def remove_duplicates_via_key(seq, key=None):
seen = set()
result = []
for item in seq:
k = key(item) if key else item
if k not in seen:
seen.add(k)
result.append(item)
return result
# Example: deduplicate a list of dictionaries by their 'id' field
data = [{'id': 1, 'val': 'a'}, {'id': 2, 'val': 'b'}, {'id': 1, 'val': 'c'}]
unique = remove_duplicates_via_key(data, key=lambda d: d['id'])
# unique → [{'id': 1, 'val': 'a'}, {'id': 2, 'val': 'b'}]
Leveraging dict.fromkeys (Python 3.7+)
Since dictionaries preserve insertion order as of Python 3.7, dict.fromkeys provides a concise one‑liner that also works for hashable items:
data = [10, 20, 10, 30, 20, 40]
unique_data = list(dict.fromkeys(data))
print(unique_data) # → [10, 20, 30, 40]
For unhashable items, you can combine this with a key function as shown earlier, or first map each element to a hashable representation, deduplicate, then map back Small thing, real impact..
Using collections.OrderedDict (for older Python versions)
If you need to support Python 3.6 or earlier, OrderedDict offers the same order‑preserving behavior:
from collections import OrderedDict
data = [10, 20, 10, 30, 20, 40]
unique_data = list(OrderedDict.fromkeys(data))
Performance Considerations
- Set‑based methods (
set,dict.fromkeys,OrderedDict.fromkeys) run in average‑case O(n) time and O(n) extra space. - Loop with a list
seendegrades to O(n²) in the worst case because eachitem not in seentriggers a linear scan; it is only practical for tiny lists or when the elements are unhashable and no suitable key exists. - Key‑based deduplication retains O(n) performance while allowing custom equivalence logic (e.g., case‑insensitive strings, ignoring whitespace).
- For very large datasets that exceed memory, consider streaming approaches: write unique items to a temporary file or database while keeping a compact hash set (or Bloom filter) of seen keys.
Quick Reference Table
| Method | Preserves Order? But | Handles Unhashable? On top of that, | Typical Complexity |
|---|---|---|---|
list(set(seq)) |
No | No (hashable only) | O(n) time, O(n) space |
Loop + set |
Yes | No | O(n) time, O(n) space |
Loop + list (linear search) |
Yes | Yes | O(n²) time, O(n) space |
dict. In real terms, fromkeys(seq) (Py≥3. 7) |
Yes | No | O(n) time, O(n) space |
| `OrderedDict. |