Getting a list of files in a directory is one of the most fundamental tasks in Python scripting, automation, and data processing workflows. In real terms, python offers several modules to achieve this, each with distinct advantages regarding readability, performance, and feature set. Understanding the differences between os.That said, whether you are building a data pipeline that ingests CSV files, writing a cleanup script to remove temporary logs, or creating a backup utility, you need a reliable way to discover what files exist on the filesystem. walk, os.So listdir, glob, os. scandir, and the modern pathlib module allows you to choose the right tool for your specific use case.
The Modern Standard: Using pathlib
Introduced in Python 3.4, the pathlib module provides an object-oriented approach to filesystem paths. It is widely considered the most "Pythonic" way to handle file system operations in modern codebases. Instead of manipulating strings, you work with Path objects that offer intuitive methods and properties.
To list files in a directory using pathlib, you typically use the iterdir() method for a non-recursive listing or rglob() / glob() for pattern matching and recursion No workaround needed..
from pathlib import Path
# Define the directory path
directory = Path('/path/to/your/directory')
# 1. List all entries (files and directories) non-recursively
print("--- Using iterdir() ---")
for entry in directory.iterdir():
print(entry.name)
# 2. List only files (filtering out directories)
print("\n--- Files only ---")
files = [entry for entry in directory.iterdir() if entry.is_file()]
for f in files:
print(f.name)
# 3. Recursive search for specific patterns (e.g., all .txt files)
print("\n--- Recursive .txt files ---")
txt_files = directory.rglob('*.txt')
for f in txt_files:
print(f)
Key advantages of pathlib:
- Cross-platform compatibility: Handles Windows backslashes and POSIX forward slashes automatically.
- Object-oriented: You get access to
.suffix,.stem,.parent,.stat(), and.read_text()directly on the object. - Readability: Code reads like English sentences (
path.is_file(),path.rglob('*.py')).
The Classic Approach: os.listdir and os.walk
Before pathlib, the os module was the standard way to interact with the operating system. It remains prevalent in legacy codebases and is perfectly functional for simple scripts. Even so, it returns strings rather than objects, meaning you must manually join paths using os.path.join to perform further operations like checking file size or modification time Turns out it matters..
Most guides skip this. Don't.
Non-Recursive Listing with os.listdir
This function returns a list containing the names of the entries in the directory. It does not return the full path, nor does it distinguish between files and folders.
import os
path = '/path/to/your/directory'
# Get all entries
entries = os.listdir(path)
print("All entries:")
for entry in entries:
# Must join path to check if it's a file
full_path = os.path.join(path, entry)
if os.path.
### Recursive Traversal with `os.walk`
If you need to traverse a directory tree recursively, `os.walk` is the traditional workhorse. It generates the file names in a directory tree by walking the tree either top-down or bottom-up. For each directory in the tree rooted at directory top (including top itself), it yields a 3-tuple `(dirpath, dirnames, filenames)`.
```python
import os
root_dir = '/path/to/your/directory'
for dirpath, dirnames, filenames in os.walk(root_dir):
print(f"Current Directory: {dirpath}")
# Print subdirectories
for dirname in dirnames:
print(f" Subdir: {dirname}")
# Print files
for filename in filenames:
# Construct full path if needed
full_path = os.And path. join(dirpath, filename)
print(f" File: {filename} (Size: {os.path.
**Performance Note:** `os.walk` calls `os.listdir` and `os.path.isdir` internally. On network filesystems or directories with thousands of entries, this can be slower than `os.scandir` because it makes separate system calls to stat each file.
## High Performance: `os.scandir` (Python 3.5+)
For performance-critical applications—such as scanning massive directories containing tens of thousands of files—`os.Even so, scandir` is significantly faster than `os. listdir` or `os.walk`. Plus, it returns an iterator of `os. DirEntry` objects instead of bare strings.
The `DirEntry` object caches the result of `os.Still, stat` (file type, size, modification time) when the operating system provides it during the directory scan. This avoids expensive subsequent system calls when you check `entry.Here's the thing — is_file()` or `entry. stat()`.
```python
import os
path = '/path/to/your/directory'
# Using scandir as a context manager ensures the iterator is closed properly
with os.scandir(path) as entries:
for entry in entries:
# is_file() uses cached info if available, very fast
if entry.is_file():
# stat() also uses cached info
info = entry.stat()
print(f"{entry.name} - Size: {info.st_size} bytes - Modified: {info.st_mtime}")
When to use os.scandir:
- Processing directories with >10,000 files.
- You need file metadata (size, dates) immediately during the scan.
- Building custom recursive walkers where you need fine-grained control over recursion logic.
Pattern Matching: The glob Module
The glob module finds all pathnames matching a specific pattern according to the rules used by the Unix shell (wildcards *, ?Still, listdir and fnmatch, but it returns full paths (relative or absolute) which is often more convenient than os. , character ranges [seq]). In real terms, it is essentially a wrapper around os. listdir.
import glob
# Non-recursive: all .py files in current directory
py_files = glob.glob('*.py')
print("Python files:", py_files)
# Non-recursive: full paths
full_paths = glob.glob('/path/to/dir/*.csv')
print("CSV files:", full_paths)
# Recursive (Python 3.5+): ** matches any files and zero or more directories
# Note: recursive=True is required for ** to work
all_logs = glob.glob('/path/to/dir/**/*.log', recursive=True)
print("All log files recursively:", all_logs)
Limitations of glob:
- It loads all results into memory at once (returns a
list), which can be problematic for massive directory trees. - It does not provide file metadata (size, permissions) without additional
os.statcalls.
Filtering and Practical Patterns
In real-world scenarios, you rarely want every file. On the flip side, you usually need to filter by extension, name pattern, size, or modification date. Here are common patterns combining the tools above.
Filtering by Extension (Case-Insensitive)
File extensions on Windows are case-insensitive, while Linux/macOS are case-sensitive. A strong check normalizes the case Easy to understand, harder to ignore. Nothing fancy..
from pathlib import Path
target_dir = Path('.gif', '.png', '.jpg', '.')
image_extensions = {'.Day to day, jpeg', '. webp', '.
# Using a set for O(1) lookup
images = [
f for f in target_dir.iterdir()
if f.is_file
```python
from pathlib import Path
target_dir = Path('.Day to day, ')
image_extensions = {'. jpg', '.Which means jpeg', '. png', '.Day to day, gif', '. webp', '.
# Using a set for O(1) lookup
images = [
f for f in target_dir.iterdir()
if f.is_file() and f.suffix.lower() in image_extensions
]
print(f"Found {len(images)} image files in {target_dir}")
Filtering by Size and Modification Time
When you need to restrict results further—say, only files larger than a certain threshold or modified within the last N days—you can attach additional predicates to the comprehension or filter loop:
import time
min_size = 100 * 1024 # 100 KB
max_age_days = 7
now = time.time()
cutoff = now - (max_age_days * 86400)
recent_large_images = [
f for f in images
if f.st_size >= min_size and f.That's why stat(). stat().
Because `Path.stat()` performs a system call for each file, consider using `os.scandir` when you already need the metadata; the `DirEntry` objects cache the results of `is_file()`, `stat()`, and related calls, making the filter essentially free after the initial scan:
```python
import os
from pathlib import Path
def scan_images(root: Path, min_size: int = 0, max_age_days: int = None):
cutoff = None
if max_age_days is not None:
cutoff = time.time() - (max_age_days * 86400)
with os.In real terms, stat() # cached, cheap
if info. Think about it: suffix. is_file():
continue
if entry.lower() not in image_extensions:
continue
info = entry.Because of that, scandir(root) as it:
for entry in it:
if not entry. st_size < min_size:
continue
if cutoff is not None and info.st_mtime < cutoff:
continue
yield Path(entry.
### Recursive Searches with `pathlib.rglob`
If you need to descend into subdirectories, `Path.rglob` offers a concise, generator‑based alternative to `glob` with the added benefit of returning `Path` objects directly:
```python
all_images = [
p for p in target_dir.rglob('*')
if p.is_file() and p.suffix.lower() in image_extensions
]
Because rglob yields results lazily, you can pipeline further filters without materializing the entire list:
def filter_by_size(paths, min_size):
for p in paths:
if p.stat().st_size >= min_size:
yield p
large_images = filter_by_size(all_images, min_size=500_000) # 0.5 MB
for img in large_images:
print(img, img.stat().
### Combining Tools for Maximum Efficiency
In performance‑critical code—such as scanning a backup volume with millions of files—you typically want:
1. **A single pass** over the directory tree (`os.scandir` or `os.walk`).
2. **Cached metadata** to avoid repeated `stat` calls.
3. **Generator‑based filtering** so memory usage stays constant regardless of tree size.
4. **Early exits** for cheap checks (extension, file‑type) before invoking more expensive operations (size, timestamps).
A compact implementation that satisfies all four points looks like this:
```python
import os, time
from pathlib import Path
def efficient_scan(root: Path,
extensions: set[str],
min_size: int = 0,
max_age_days: int | None = None):
cutoff = None
if max_age_days is not
```python
if max_age_days is not None:
cutoff = time.time() - (max_age_days * 86400)
with os.Consider this: scandir(root) as it:
for entry in it:
if not entry. On top of that, is_file():
continue
if entry. Because of that, suffix. Consider this: lower() not in image_extensions:
continue
try:
stat_info = entry. stat()
if stat_info.st_size < min_size:
continue
if cutoff is not None and stat_info.In real terms, st_mtime < cutoff:
continue
yield Path(entry. path)
except OSError:
# Skip files we cannot access (permission errors, broken symlinks, etc.
This final version adds a safety net by catching `OSError` exceptions that may arise when traversing deeply nested directories where some entries lack readable metadata due to permission restrictions or filesystem quirks. By swallowing these errors and moving on, the scanner remains reliable against edge cases common in production environments.
When scaling beyond a few hundred megabytes of data, the single‑pass approach described above still hits a hard limit: Python’s default recursion depth and the overhead of creating many small generator frames become noticeable. In those scenarios, two complementary strategies often improve throughput:
* **Parallel chunking** – Split the top‑level directory into sub‑ranges based on numeric prefixes (e.g., `img_001.jpg`, `img_002.jpg`) and dispatch independent workers to each range. Because each worker only walks its own slice, contention is eliminated and CPU cores are utilized fully. Libraries such as `concurrent.futures.ThreadPoolExecutor` or the native `multiprocessing.Pool` make this straightforward.
* **Memory‑mapped reads** – For extremely large image collections stored on fast storage (SSD/NVMe), `mmap()` can read file contents without loading them entirely into RAM. While not necessary for pure metadata filtering, combining `mmap` with `os.stat` eliminates the need to call `stat()` repeatedly, reducing system‑call overhead dramatically.
Both techniques preserve the core philosophy introduced at the start of the article: never perform unnecessary work. Keep the heavy lifting (reading actual pixel data) until after you have identified candidate files. If your downstream task involves per‑file computation—executing AI inference, extracting EXIF data, or transcoding—the scanned list should be fed straight into that consumer, leaving the raw image bytes untouched until they are truly needed.
---
## Summary
Efficiently locating images across a directory tree hinges on four design principles:
1. **Single pass** – Traverse the tree once, using either `os.scandir` (or `os.walk` for recursive walks) rather than building intermediate lists.
2. **Cached metadata** – Rely on `DirEntry` objects, which store the result of `is_file()`, `stat()` and their related helpers after the first lookup, turning subsequent queries into cheap lookups.
3. **Lazy generation** – Yield `Path` objects one at a time instead of materialising the whole collection, guaranteeing constant memory consumption even for multi‑gigabyte datasets.
4. **Early pruning** – Apply inexpensive filters (suffix check, minimum size, age limits) before any costly operation (full `stat()`, time‑stamp comparison). This reduces the number of expensive syscalls and keeps the hot path short.
By adhering to these guidelines, you obtain a scanner that scales gracefully from a handful of gigabytes to several terabytes while staying responsive to interactive use. When combined with optional parallelism and memory‑mapping for bulk workloads, the pattern becomes a cornerstone of high‑per
When combined with optional parallelism and memory‑mapping for bulk workloads, the pattern becomes a cornerstone of high‑performance data pipelines. The scanner can be plugged directly into downstream consumers—whether they are AI inference engines, EXIF extractors, or video transcoders—without ever loading image bytes into memory prematurely. Below are a few concrete ways to tie everything together.
### 1. Feeding a ThreadPoolExecutor
```python
from concurrent.futures import ThreadPoolExecutor
from pathlib import Path
def process_image(path: Path):
# Heavy‑weight operation: read pixels, run inference, etc.
In real terms, open('rb') as f:
# Example: a dummy operation that mimics inference
data = f. with path.read(1024) # read only what you need
return some_model.
def bulk_process(root: Path, pattern: str = '*.jpg'):
# The scanner yields Path objects lazily
candidates = image_scanner(root, suffix=pattern)
# Dispatch work to a pool of workers
with ThreadPoolExecutor(max_workers=8) as exe:
results = exe.map(process_image, candidates)
for res in results:
yield res
The image_scanner function (the one described in the earlier sections) guarantees that only files passing the cheap filters reach the executor. So naturally, the worker threads spend their time on the real work, not on filesystem syscalls And that's really what it comes down to..
2. Memory‑Mapped Bulk Processing
When the storage medium can sustain the bandwidth of mmap, you can avoid the read() call inside process_image altogether:
import mmap
import os
def mmap_process(path: Path):
with open(path, 'rb') as f:
with mmap.fileno(), 0, access=mmap.That's why mmap(f. ACCESS_READ) as mm:
# mm behaves like a bytes‑like object; you can slice it directly
return some_model.
Because the scanner already performed an `os., > 10 MiB). That's why stat` (or used `DirEntry. st_size`), you know the file size and can decide whether mapping is worthwhile (e.For tiny thumbnails, a regular `open().Practically speaking, g. read()` is often faster.
### 3. Chaining Filters for Domain‑Specific Needs
Real‑world pipelines rarely stop at a simple suffix check. The scanner’s early‑pruning design makes it trivial to add more inexpensive predicates:
```python
def is_recent(path: Path, days: int = 30) -> bool:
# DirEntry already holds st_mtime; no extra stat() needed
return (time.time() - path.stat().st_mtime) < days * 86400
def size_range(path: Path, min_bytes: int, max_bytes: int) -> bool:
sz = path.stat().st_size
return min_bytes <= sz <= max_bytes
candidates = image_scanner(
root,
suffix='*.png',
predicates=[
lambda p: is_recent(p, days=7),
lambda p: size_range(p, 1024, 10_000_000),
],
)
All predicates receive a Path that already carries the cached DirEntry information, so each check is essentially a dictionary lookup.
4. Integration with Asynchronous Workflows
If your application is async (e.g., a web service that ingests uploads), you can run the scanner in a thread pool to avoid blocking the event loop:
import asyncio
from concurrent.futures import ThreadPoolExecutor
async def async_scan(root: Path):
loop = asyncio.And get_running_loop()
with ThreadPoolExecutor() as pool:
candidates = await loop. run_in_executor(
pool, image_scanner, root
)
# candidates is a generator; you can iterate it asynchronously
async for path in async_generator(candidates):
# schedule heavy work
await loop.
The scanner’s lazy nature ensures that the async task never accumulates a huge list of paths in memory.
---
## Conclusion
The article has outlined a pragmatic, four‑principle approach to scanning large directories of images:
1. **Single pass** traversal using `os.scandir`/`Path.rglob` to avoid intermediate collections.
2. **Cached metadata** via `DirEntry` objects, turning repeated `stat()` calls into cheap lookups.
3. **Lazy generation** of `Path` objects, keeping memory footprints constant even for terabyte‑scale datasets.
4