How To Remove Duplicates From A List In Python

4 min read

Removing duplicates from a list in Python is one of the most common data-cleaning tasks developers face when working with real-world data. But whether you are processing user input, preparing data for analysis, or building a clean dataset for machine learning, duplicate values can cause incorrect counts, skewed results, and inefficient memory usage. This article explains practical ways to remove duplicates from a list in Python, including methods that preserve order, handle unhashable items, and improve performance No workaround needed..

Introduction

In Python, a list is an ordered, mutable collection that can contain repeated values. For example:

data = [10, 20, 10, 30, 20, 40]

The list above contains duplicate values: 10 and 20 appear more than once. If your program needs unique values only, you need a reliable way to remove

Removing Duplicates with a Set

The simplest approach to eliminate duplicates is to convert the list to a set, which inherently contains only unique elements. That said, this method does not preserve the original order of elements:

data = [10, 20, 10, 30, 20, 40]
unique_data = list(set(data))
print(unique_data)  # Output: [10, 20, 30, 40] or any order

While effective for small datasets, this method is unsuitable when order matters. For ordered deduplication, alternative strategies are required And it works..

Preserving Order with a Loop

To maintain the original order while removing duplicates, iterate through the list and track seen elements with a helper set:

def remove_duplicates_ordered(input_list):
    seen = set()
    result = []
    for item in input_list:
        if item not in seen:
            seen.add(item)
            result.append(item)
    return result

data = [10, 20, 10, 30, 20, 40]
print(remove_duplicates_ordered(data))  # Output: [10, 20, 30, 40]

This approach ensures uniqueness while preserving insertion order, making it ideal for scenarios like cleaning user

…cleaning user input, maintaining audit trails, or preparing time‑series data where the chronological order of events must stay intact.

Handling Unhashable Items

When the list contains mutable objects such as dictionaries, lists, or sets, the simple seen = set() trick fails because these types cannot be hashed. A common workaround is to replace the hash‑based seen with a list (or another suitable container) and rely on equality checks:

def remove_duplicates_unhashable(seq):
    seen = []
    result = []
    for item in seq:
        if item not in seen:          # linear search; acceptable for small‑to‑moderate sizes
            seen.append(item)
            result.append(item)
    return result

If the unhashable elements are themselves hashable after a transformation (e.g., converting inner lists to tuples), you can keep the fast set lookup:

def remove_duplicates_via_key(seq, key=None):
    seen = set()
    result = []
    for item in seq:
        k = key(item) if key else item
        if k not in seen:
            seen.add(k)
            result.append(item)
    return result

# Example: deduplicate a list of dictionaries by their 'id' field
data = [{'id': 1, 'val': 'a'}, {'id': 2, 'val': 'b'}, {'id': 1, 'val': 'c'}]
unique = remove_duplicates_via_key(data, key=lambda d: d['id'])
# unique → [{'id': 1, 'val': 'a'}, {'id': 2, 'val': 'b'}]

Leveraging dict.fromkeys (Python 3.7+)

Since dictionaries preserve insertion order as of Python 3.7, dict.fromkeys provides a concise one‑liner that also works for hashable items:

data = [10, 20, 10, 30, 20, 40]
unique_data = list(dict.fromkeys(data))
print(unique_data)   # → [10, 20, 30, 40]

For unhashable items, you can combine this with a key function as shown earlier, or first map each element to a hashable representation, deduplicate, then map back Small thing, real impact..

Using collections.OrderedDict (for older Python versions)

If you need to support Python 3.6 or earlier, OrderedDict offers the same order‑preserving behavior:

from collections import OrderedDict

data = [10, 20, 10, 30, 20, 40]
unique_data = list(OrderedDict.fromkeys(data))

Performance Considerations

  • Set‑based methods (set, dict.fromkeys, OrderedDict.fromkeys) run in average‑case O(n) time and O(n) extra space.
  • Loop with a list seen degrades to O(n²) in the worst case because each item not in seen triggers a linear scan; it is only practical for tiny lists or when the elements are unhashable and no suitable key exists.
  • Key‑based deduplication retains O(n) performance while allowing custom equivalence logic (e.g., case‑insensitive strings, ignoring whitespace).
  • For very large datasets that exceed memory, consider streaming approaches: write unique items to a temporary file or database while keeping a compact hash set (or Bloom filter) of seen keys.

Quick Reference Table

Method Preserves Order? But Handles Unhashable? On top of that, Typical Complexity
list(set(seq)) No No (hashable only) O(n) time, O(n) space
Loop + set Yes No O(n) time, O(n) space
Loop + list (linear search) Yes Yes O(n²) time, O(n) space
dict. In real terms, fromkeys(seq) (Py≥3. 7) Yes No O(n) time, O(n) space
`OrderedDict.
What's New

This Week's Picks

In That Vein

Related Reading

Thank you for reading about How To Remove Duplicates From A List In Python. We hope the information has been useful. Feel free to contact us if you have any questions. See you next time — don't forget to bookmark!
⌂ Back to Home