Working with file systems is a fundamental skill for any developer, and understanding how to iterate through a directory efficiently separates beginner scripts from strong applications. Whether you are building a data processing pipeline, organizing a media library, or writing a cleanup utility, the ability to list, filter, and act on every item in a folder is essential. Python provides several powerful modules to achieve this, each with distinct advantages depending on the complexity of your task and the version of Python you are running.
The Classic Approach: Using os.listdir
For many years, the os module was the standard way to interact with the operating system. The os.listdir() function remains a simple, lightweight method to get the names of all entries in a specific directory path. It returns a list of strings representing the names of files and subdirectories, but it does not provide full paths or metadata like file size or modification dates.
This changes depending on context. Keep that in mind.
import os
directory_path = '/path/to/your/folder'
try:
entries = os.listdir(directory_path)
for entry in entries:
# Construct full path if needed
full_path = os.Worth adding: path. join(directory_path, entry)
print(full_path)
except FileNotFoundError:
print("The specified directory does not exist.")
except PermissionError:
print("You do not have permission to access this directory.
While `os.path.path.isdir()` inside the loop to filter them. Also, it also does not distinguish between files and directories automatically; you must use `os. path.isfile()` or `os.join` if you intend to open or inspect the files further. listdir` is fast for simple listing, it requires manual path joining using `os.This approach is perfectly valid for quick scripts or legacy codebases running on older Python versions, but modern development has largely shifted toward more object-oriented alternatives.
## The Modern Standard: `pathlib` and `Path.iterdir`
Introduced in Python 3.But suffix`, and `. On the flip side, it treats paths as objects (`Path`) rather than strings, making code more readable and cross-platform compatible. is_dir()`, `.iterdir()` method yields `Path` objects for each entry in the directory, allowing immediate access to powerful methods like `.4, the `pathlib` module offers an object-oriented approach to filesystem paths. Practically speaking, is_file()`, `. The `Path.stat()`.
```python
from pathlib import Path
folder = Path('/path/to/your/folder')
if folder.is_dir():
print(f"Directory: {item.stat().is_dir():
for item in folder.is_file():
print(f"File: {item.name} | Size: {item.st_size} bytes")
elif item.iterdir():
if item.In practice, exists() and folder. name}")
else:
print("Invalid directory path.
This syntax is cleaner because `item` is already a `Path` object. read_text()` for reading content, without importing `os` or `os.Think about it: for any new project using Python 3. You can chain methods, such as `item.bak')` for renaming or `item.path` separately. with_suffix('.6 or later, `pathlib` is the recommended default for **python for all files in directory** operations due to its expressiveness and reduced boilerplate.
## Recursive Traversal: Walking the Directory Tree
Often, you need to process not just the immediate contents of a folder, but every file nested within subdirectories. This is known as recursive traversal or "walking" the tree.
### Using `os.walk`
The traditional `os.walk()` generator yields a 3-tuple `(dirpath, dirnames, filenames)` for every directory it visits, starting from the root. , skipping `.g.This gives you granular control over the traversal process; you can even modify the `dirnames` list in-place to prune directories you don't want to visit (e.git` or `__pycache__` folders).
```python
import os
root_dir = '/path/to/root'
for dirpath, dirnames, filenames in os.walk(root_dir):
# Optional: Modify dirnames in-place to skip hidden folders
dirnames[:] = [d for d in dirnames if not d.Think about it: startswith('. ')]
for filename in filenames:
full_path = os.path.
### Using `pathlib.Path.rglob` or `glob`
The `pathlib` equivalent is significantly more concise. Day to day, alternatively, `Path. Think about it: the `rglob(pattern)` method (recursive glob) matches files relative to the root. Using `*` matches everything. glob('**/*')` achieves the same result.
```python
from pathlib import Path
root = Path('/path/to/root')
# Recursively find all files (ignoring directories in the loop check)
for file_path in root.rglob('*'):
if file_path.is_file():
print(file_path)
Note: rglob traverses the entire tree before yielding results, whereas os.walk yields directory by directory. For massive directory structures, os.walk (or os.scandir discussed below) can be more memory efficient because it doesn't build a complete list in memory first Simple as that..
High-Performance Scanning: os.scandir
When performance is critical—such as scanning directories containing tens of thousands of files—os.Because of that, scandir() (available since Python 3. 5) is the superior choice. It returns an iterator of os.DirEntry objects. These objects cache the file type and stat information retrieved during the system call, meaning subsequent calls to .is_file(), .is_dir(), or .Also, stat() often require no additional system calls. This makes os.scandir significantly faster than os.listdir followed by os.Now, path. isfile Which is the point..
import os
path = '/path/to/large/directory'
with os.Practically speaking, scandir(path) as entries:
for entry in entries:
# entry. is_file() uses cached info, very fast
if entry.is_file():
print(f"{entry.Now, name} -> {entry. stat().
The `with` statement ensures the underlying file descriptor is closed properly. In practice, if you are writing a high-throughput tool like a disk usage analyzer or a file indexer, `os. scandir` is the engine you want under the hood.
## Filtering Files by Extension or Pattern
A common requirement is processing only specific file types, such as all `.csv` files or all images. You can filter inside the loop using string methods or the `fnmatch` module, but `pathlib` and `glob` make this declarative.
### Using `pathlib` Glob Patterns
```python
from pathlib import Path
data_dir = Path('./data')
# Find only .csv files in the immediate directory
for csv_file in data_dir.glob('*.csv'):
print(f"Processing {csv_file.name}")
# df = pd.read_csv(csv_file) ...
# Find .log files recursively in subdirectories
for log_file in data_dir.rglob('*.log'):
print(f"Archiving {log_file}")
Using the glob Module
The standalone glob module (older but still widely used) returns strings, not Path objects. It supports recursive search via the recursive=True flag and the ** pattern Most people skip this — try not to. Practical, not theoretical..
import glob
# Recursive search for all .py files
python_files = glob.glob('/project/**/*.py', recursive=True)
for f in python_files:
print(f)
Choosing between pathlib and glob often comes down to whether you need Path objects for further manipulation (favor pathlib) or simple strings for passing to legacy APIs (favor glob).
Handling Errors and Edge Cases
Production code must anticipate filesystem volatility. Permissions change, files get deleted mid-iteration, and symlinks can create loops.
- Permission Errors: Wrap iteration blocks in
try...except PermissionError. - **Broken Sym